# Vichar — Full Documentation
> Vichar is an open-source, OpenAI-compatible API gateway for routing, managing, and analyzing requests across LLM providers. Use one API key, track usage and cost, configure caching and guardrails, and self-host or use the managed cloud. Current models and pricing: https://app.vichar.io/models
API base URL: https://api.vichar.io/v1 · Docs: https://docs.vichar.io · Site: https://app.vichar.io
This file concatenates the full text of every documentation page below.
# Introduction to Vichar
URL: https://docs.vichar.io/
Vichar is an open-source API gateway that sits between your applications and LLM providers like OpenAI, Anthropic, Google AI Studio, and more. It provides a unified, OpenAI-compatible API interface with built-in cost tracking, caching, and intelligent routing.
## Products [#products]
These docs cover four products. Pick one from the menu at the top of the sidebar, or start here:
## How it works [#how-it-works]
Point your existing SDK at `https://api.vichar.io/v1` (or your self-hosted instance) and authenticate with an Vichar API key — no code rewrites. The gateway routes each request to the right provider, tracks tokens and cost per model, provider, project, and API key, and fails over to a healthy provider when one errors. Pay per-token with prepaid credits at provider list rates.
## Why use a gateway? [#why-use-a-gateway]
* **One integration** — switch models or providers by changing a model string, not your code.
* **Cost visibility** — usage analytics and cost breakdowns across every provider in one dashboard.
* **Reliability** — automatic provider failover when a provider fails, plus opt-in response caching per project.
* **No lock-in** — open source (AGPLv3), self-hostable, and OpenAI-compatible end to end.
## Take the tour [#take-the-tour]
## Features [#features]
All features are documented under https://docs.vichar.io/features; each feature page is included in full in this file.
## AI Tooling [#ai-tooling]
Vichar is built to work seamlessly with AI agents and development tools.
AI tooling: https://docs.vichar.io/llms.txt (docs index for LLMs), https://docs.vichar.io/llms-full.txt (this file), and https://docs.vichar.io/developers/mcp (MCP server).
## Next Steps [#next-steps]
* [**Quickstart**](https://docs.vichar.io/quick-start) — Get up and running in minutes
* [**Overview**](https://docs.vichar.io/overview) — Learn more about what Vichar offers
* [**Self-Host**](https://docs.vichar.io/self-host) — Deploy on your own infrastructure
# Overview
URL: https://docs.vichar.io/overview
Vichar is an open-source API gateway for Large Language Models (LLMs). It acts as a middleware between your applications and various LLM providers, allowing you to:
* Route requests to multiple LLM providers (OpenAI, Anthropic, Google AI Studio, and others)
* Manage API keys for different providers in one place
* Track token usage and costs across all your LLM interactions
* Analyze performance metrics to optimize your LLM usage
## Analyzing Your LLM Requests [#analyzing-your-llm-requests]
Vichar provides detailed insights into your LLM usage:
* **Usage Metrics**: Track the number of requests, tokens used, and response times
* **Cost Analysis**: Monitor spending across different models and providers
* **Performance Tracking**: Identify patterns and optimize your prompts based on actual usage data
* **Breakdown by Model**: Compare different models' performance and cost-effectiveness
All this data is automatically collected and presented in an intuitive dashboard, helping you make informed decisions about your LLM strategy.
## Getting Started [#getting-started]
Using Vichar is simple. Just swap out your current LLM provider URL with the Vichar API endpoint:
```bash
curl -X POST https://api.vichar.io/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
]
}'
```
Vichar maintains compatibility with the OpenAI API format, making migration seamless. Note that unknown or unsupported request parameters (for example `stop`, `seed`, `logprobs`, or `logit_bias`) are accepted and silently ignored rather than rejected, so a request carrying them still succeeds — the parameter just has no effect.
## Hosted vs. Self-Hosted [#hosted-vs-self-hosted]
You can use Vichar in two ways:
* **Hosted Version**: For immediate use without setup, visit [app.vichar.io](https://app.vichar.io) to create an account and get an API key.
* **Self-Hosted**: Deploy Vichar on your own infrastructure for complete control over your data and configuration.
The self-hosted version offers additional customization options and ensures your LLM traffic never leaves your infrastructure if desired.
# Quickstart
URL: https://docs.vichar.io/quick-start
Welcome to **Vichar**—a single drop‑in endpoint that lets you call today’s best large‑language models while keeping **your existing code** and development workflow intact.
> **TL;DR** — Point your HTTP requests to `https://api.vichar.io/v1/…`, supply your `LLM_GATEWAY_API_KEY`, and you’re done.
***
## 1 · Get an API key [#1--get-an-api-key]
1. Sign in to the dashboard.
2. Create a new Project → *Copy the key*.
3. Export it in your shell (or a `.env` file):
```bash
export LLM_GATEWAY_API_KEY="vichar_XXXXXXXXXXXXXXXX"
```
***
## 2 · Pick your language [#2--pick-your-language]
```bash
curl -X POST https://api.vichar.io/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
]
}'
```
```typescript
const response = await fetch("https://api.vichar.io/v1/chat/completions", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
},
body: JSON.stringify({
model: "gpt-4o",
messages: [{ role: "user", content: "Hello, how are you?" }],
}),
});
if (!response.ok) {
throw new Error(`HTTP error! status: ${response.status}`);
}
const data = await response.json();
console.log(data.choices[0].message.content);
```
```tsx
import { useState } from "react";
function ChatComponent() {
const [response, setResponse] = useState("");
const [loading, setLoading] = useState(false);
const sendMessage = async () => {
setLoading(true);
try {
const res = await fetch("https://api.vichar.io/v1/chat/completions", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.REACT_APP_LLM_GATEWAY_API_KEY}`,
},
body: JSON.stringify({
model: "gpt-4o",
messages: [{ role: "user", content: "Hello, how are you?" }],
}),
});
if (!res.ok) {
throw new Error(`HTTP error! status: ${res.status}`);
}
const data = await res.json();
setResponse(data.choices[0].message.content);
} catch (error) {
console.error("Error:", error);
} finally {
setLoading(false);
}
};
return (
{response &&
{response}
}
);
}
export default ChatComponent;
```
```typescript
// app/api/chat/route.ts
import { NextRequest, NextResponse } from "next/server";
export async function POST(request: NextRequest) {
const { message } = await request.json();
const response = await fetch("https://api.vichar.io/v1/chat/completions", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
},
body: JSON.stringify({
model: "gpt-4o",
messages: [{ role: "user", content: message }],
}),
});
if (!response.ok) {
return NextResponse.json(
{ error: "Failed to get response" },
{ status: response.status },
);
}
const data = await response.json();
return NextResponse.json({
message: data.choices[0].message.content,
});
}
// Usage in component:
// const response = await fetch('/api/chat', {
// method: 'POST',
// headers: { 'Content-Type': 'application/json' },
// body: JSON.stringify({ message: 'Hello, how are you?' })
// });
```
```python
import requests
import os
response = requests.post(
'https://api.vichar.io/v1/chat/completions',
headers={
'Content-Type': 'application/json',
'Authorization': f'Bearer {os.getenv("LLM_GATEWAY_API_KEY")}'
},
json={
'model': 'gpt-4o',
'messages': [
{'role': 'user', 'content': 'Hello, how are you?'}
]
}
)
response.raise_for_status()
print(response.json()['choices'][0]['message']['content'])
```
```java
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.net.URI;
String apiKey = System.getenv("LLM_GATEWAY_API_KEY");
String requestBody = """
{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
]
}
""";
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.vichar.io/v1/chat/completions"))
.header("Content-Type", "application/json")
.header("Authorization", "Bearer " + apiKey)
.POST(HttpRequest.BodyPublishers.ofString(requestBody))
.build();
HttpResponse response = HttpClient.newHttpClient()
.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
```
```rust
use reqwest::Client;
use serde_json::json;
use std::env;
#[tokio::main]
async fn main() -> Result<(), Box> {
let client = Client::new();
let api_key = env::var("LLM_GATEWAY_API_KEY")?;
let response = client
.post("https://api.vichar.io/v1/chat/completions")
.header("Content-Type", "application/json")
.header("Authorization", format!("Bearer {}", api_key))
.json(&json!({
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
]
}))
.send()
.await?;
let result: serde_json::Value = response.json().await?;
println!("{}", result["choices"][0]["message"]["content"]);
Ok(())
}
```
```go
package main
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"os"
)
type ChatRequest struct {
Model string `json:"model"`
Messages []Message `json:"messages"`
}
type Message struct {
Role string `json:"role"`
Content string `json:"content"`
}
func main() {
apiKey := os.Getenv("LLM_GATEWAY_API_KEY")
requestBody := ChatRequest{
Model: "gpt-4o",
Messages: []Message{{Role: "user", Content: "Hello, how are you?"}},
}
jsonData, _ := json.Marshal(requestBody)
req, _ := http.NewRequest("POST", "https://api.vichar.io/v1/chat/completions", bytes.NewBuffer(jsonData))
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Authorization", "Bearer "+apiKey)
client := &http.Client{}
resp, _ := client.Do(req)
defer resp.Body.Close()
fmt.Println("Response received")
}
```
```php
'gpt-4o',
'messages' => [
['role' => 'user', 'content' => 'Hello, how are you?']
]
];
$options = [
'http' => [
'header' => [
'Content-Type: application/json',
'Authorization: Bearer ' . $apiKey
],
'method' => 'POST',
'content' => json_encode($data)
]
];
$context = stream_context_create($options);
$response = file_get_contents(
'https://api.vichar.io/v1/chat/completions',
false,
$context
);
if ($response === FALSE) {
throw new Exception('Request failed');
}
$result = json_decode($response, true);
echo $result['choices'][0]['message']['content'];
?>
```
```ruby
require 'net/http'
require 'json'
require 'uri'
uri = URI('https://api.vichar.io/v1/chat/completions')
http = Net::HTTP.new(uri.host, uri.port)
http.use_ssl = true
request = Net::HTTP::Post.new(uri)
request['Content-Type'] = 'application/json'
request['Authorization'] = "Bearer #{ENV['LLM_GATEWAY_API_KEY']}"
request.body = {
model: 'gpt-4o',
messages: [
{ role: 'user', content: 'Hello, how are you?' }
]
}.to_json
response = http.request(request)
if response.code != '200'
raise "HTTP Error: #{response.code}"
end
result = JSON.parse(response.body)
puts result['choices'][0]['message']['content']
```
***
## 3 · SDK integrations [#3--sdk-integrations]
```ts title="ai-sdk.ts"
import { llmgateway } from "@llmgateway/ai-sdk-provider";
import { generateText } from "ai";
const { text } = await generateText({
model: llmgateway("gpt-4o"),
prompt: "Write a vegetarian lasagna recipe for 4 people.",
});
```
```ts title="vercel-ai-sdk.ts"
import { createOpenAI } from "@ai-sdk/openai";
const llmgateway = createOpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY!,
});
const completion = await llmgateway.chat({
model: "gpt-4o",
messages: [{ role: "user", content: "Hello, how are you?" }],
});
console.log(completion.choices[0].message.content);
```
```ts title="openai-sdk.ts"
import OpenAI from "openai";
const openai = new OpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const completion = await openai.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Hello, how are you?" }],
});
console.log(completion.choices[0].message.content);
```
***
## 4 · Going further [#4--going-further]
* **Streaming**: pass `stream: true` to any request—the gateway normalizes every provider's stream into OpenAI-format SSE chunks, with a final chunk carrying `usage` and routing metadata before `data: [DONE]`.
* **Monitoring**: Every call appears in the dashboard with latency, cost & provider breakdown.
***
## 5 · FAQ [#5--faq]
See the [Models page](https://app.vichar.io/dashboard).
Unlike OpenRouter, we offer:
Full self-hosting capabilities, giving you complete control over your
infrastructure
Enhanced analytics with deeper insights into your model usage and
performance
Our pricing structure is designed to be flexible and cost-effective: See the
[Pricing section](https://app.vichar.io#pricing).
***
## 6 · Next steps [#6--next-steps]
* Read [Self host docs](https://docs.vichar.io/self-host) guide.
* Drop into our [GitHub](https://github.com/vicharai/api) for help or feature requests.
Happy building! ✨
# Health check
URL: https://docs.vichar.io/health
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Create speech
URL: https://docs.vichar.io/v1_audio_speech
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Create transcription
URL: https://docs.vichar.io/v1_audio_transcriptions
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Chat Completions
URL: https://docs.vichar.io/v1_chat_completions
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Embeddings
URL: https://docs.vichar.io/v1_embeddings
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Edit image
URL: https://docs.vichar.io/v1_images_edits
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Create image
URL: https://docs.vichar.io/v1_images_generations
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Retrieve key status
URL: https://docs.vichar.io/v1_key_retrieve
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Anthropic Messages
URL: https://docs.vichar.io/v1_messages
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Models
URL: https://docs.vichar.io/v1_models
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Moderations
URL: https://docs.vichar.io/v1_moderations
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# OCR
URL: https://docs.vichar.io/v1_ocr
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Create a realtime client secret
URL: https://docs.vichar.io/v1_realtime_client_secrets
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Rerank
URL: https://docs.vichar.io/v1_rerank
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# System One
URL: https://docs.vichar.io/v1_systemone
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Video content
URL: https://docs.vichar.io/v1_videos_content
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Create video
URL: https://docs.vichar.io/v1_videos_create
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Video log content
URL: https://docs.vichar.io/v1_videos_log_content
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Retrieve video
URL: https://docs.vichar.io/v1_videos_retrieve
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# AI SDK Gateway protocol
URL: https://docs.vichar.io/developers/ai-sdk-gateway-protocol
When you pass a bare model string to the AI SDK — `streamText({ model: "anthropic/claude-sonnet-5" })` — the SDK resolves it through its **default provider**, `@ai-sdk/gateway`. That provider does not speak the OpenAI Chat Completions format: it has its own wire protocol, `LanguageModelV*CallOptions` in and `LanguageModelV*` parts out.
Vichar implements that protocol, so an app written against the Vercel AI Gateway runs here with **no code change** — only its base URL and API key are repointed.
## Repoint the default provider [#repoint-the-default-provider]
```ts
import { createGateway } from "@ai-sdk/gateway";
globalThis.AI_SDK_DEFAULT_PROVIDER = createGateway({
baseURL: "https://api.vichar.io/v4/ai",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
```
Put this wherever your app runs before its first model call — a Next.js [`instrumentation.ts`](https://nextjs.org/docs/app/guides/instrumentation), a server entrypoint, or a platform-injected preamble. Every bare model string in the app then resolves through Vichar.
`@ai-sdk/gateway` is already a transitive dependency of `ai`, so there is nothing extra to install.
`@ai-sdk/gateway` reads its API key from `AI_GATEWAY_API_KEY` but has **no**
environment variable for the base URL — it is a constructor option only. That
is why repointing takes this one line rather than an env var.
Or construct the provider explicitly and pass it per call:
```ts
import { createGateway } from "@ai-sdk/gateway";
import { streamText } from "ai";
const gateway = createGateway({
baseURL: "https://api.vichar.io/v4/ai",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const result = streamText({
model: gateway("anthropic/claude-sonnet-5"),
prompt: "Hello!",
});
```
## Base URL per AI SDK version [#base-url-per-ai-sdk-version]
The protocol carries the language-model specification version in a request header, and every prefix below serves the same surface — pick the one matching the `@ai-sdk/gateway` your app has, so the default path stays intact:
| AI SDK | Spec version | Base URL |
| ------ | ------------ | ----------------------------- |
| 5 | 2 | `https://api.vichar.io/v1/ai` |
| 6 | 3 | `https://api.vichar.io/v3/ai` |
| 7 | 4 | `https://api.vichar.io/v4/ai` |
## Model IDs [#model-ids]
Model IDs use the provider-pinned `provider/model` form (`anthropic/claude-sonnet-5`, `openai/gpt-4o`) — the same convention AI Gateway IDs use, so existing model strings resolve unchanged.
Vichar's smart-routing IDs work here too: pass a bare model ID (`gpt-4o`) to let the gateway pick a provider, or `auto` to let it pick the model. Those are not returned by `getAvailableModels()` because they do not name one provider, but they are accepted.
## Listing models [#listing-models]
```ts
const { models } = await gateway.getAvailableModels();
```
Returns one entry per active provider mapping, with pricing and an AI SDK `specification` block. This is what backs a model picker built on `GatewayModel[]`.
## Credits [#credits]
```ts
const { balance, totalUsed } = await gateway.getCredits();
```
`balance` is the organization's remaining credit balance and `totalUsed` its lifetime credits spend.
## Gateway-only options [#gateway-only-options]
Features that have no field in the AI SDK call options — reasoning effort, service tier, routing strategy, prompt cache keys, plugins — are set through the `llmgateway` provider options namespace, which is passed onto the underlying request:
```ts
const result = streamText({
model: gateway("openai/gpt-5.6-terra"),
prompt: "Hello!",
providerOptions: {
llmgateway: {
reasoning_effort: "high",
routing: "price",
},
},
});
```
Any field the [chat completions API](https://docs.vichar.io/api-reference) accepts works here, except `model`, `messages` and `stream`, which this surface owns.
## Web search [#web-search]
The provider-native web search tools serialize to provider-defined tools, and the gateway maps them onto its [native web search](https://docs.vichar.io/features/web-search):
```ts
import { openai } from "@ai-sdk/openai";
const result = streamText({
model: gateway("openai/gpt-4o"),
prompt: "What happened in the news today?",
tools: { web_search: openai.tools.webSearch() },
});
```
Recognised tools: `openai.web_search`, `openai.web_search_preview`, `anthropic.web_search_20250305`, `anthropic.web_search_20260209`, `google.google_search`. Search results come back as `source-url` message parts plus a provider-executed tool call under the name you bound the tool to, so the AI SDK's sources UI works unchanged.
A provider-defined tool the gateway cannot map is reported as an `unsupported-tool` warning on the result rather than failing the request.
## What is not covered [#what-is-not-covered]
This surface implements language models. Embeddings, images, video, speech, transcription, reranking and realtime are served by the [OpenAI-compatible endpoints](https://docs.vichar.io/api-reference) — use [`@llmgateway/ai-sdk-provider`](https://docs.vichar.io/developers/ai-sdk) for those.
These call options have no chat completions equivalent and are reported as `unsupported` warnings: `stopSequences`, `seed`, `topK`.
# Image Generation with the AI SDK
URL: https://docs.vichar.io/developers/ai-sdk-images
The `@llmgateway/ai-sdk-provider` package supports image generation both through the AI SDK's dedicated `generateImage` function and through chat-based image models that stream images as part of a conversation.
## generateImage [#generateimage]
Use `llmgateway.image()` to get an image model:
```typescript
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { generateImage } from "ai";
import { writeFileSync } from "fs";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const result = await generateImage({
model: llmgateway.image("gemini-3-pro-image"),
prompt:
"A cozy cabin in a snowy mountain landscape at night with aurora borealis",
size: "1024x1024",
n: 1,
// aspectRatio and quality are model-specific — only some providers honor them.
// aspectRatio works on Gemini image models; OpenAI gpt-image-2 ignores it
// (use a literal WxH `size` instead).
aspectRatio: "16:9",
// quality works on OpenAI gpt-image-2 ("low" | "medium" | "high" | "auto")
// and moderation ("auto" | "low") on GPT Image models.
// The AI SDK only forwards these through providerOptions.
providerOptions: {
llmgateway: { quality: "high", moderation: "low" },
},
});
result.images.forEach((image, i) => {
const buf = Buffer.from(image.base64, "base64");
writeFileSync(`image-${i}.png`, buf);
});
```
Which sizes, aspect ratios, and quality settings a model accepts depends on
the model — see [Image Generation](https://docs.vichar.io/features/image-generation) for the full
parameter reference and per-model behavior.
## Chat-based image models [#chat-based-image-models]
Multimodal models like `gemini-3-pro-image` can return images inside a chat conversation. Use `llmgateway.chat()` with `streamText` in a route handler:
```typescript
// app/api/chat/route.ts
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { convertToModelMessages, streamText } from "ai";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
export async function POST(req: Request) {
const { messages } = await req.json();
const result = streamText({
model: llmgateway.chat("gemini-3-pro-image"),
messages: convertToModelMessages(messages),
});
return result.toUIMessageStreamResponse();
}
```
On the client, image parts arrive as message file parts that you can render with the AI Elements `Image` component or a plain `` tag with a data URL. See [Image Generation](https://docs.vichar.io/features/image-generation) for the complete `useChat` frontend example.
## Video and audio [#video-and-audio]
The AI SDK does not yet cover the gateway's video and speech endpoints — call them over REST instead:
* [Video Generation](https://docs.vichar.io/features/video-generation) — `POST /v1/videos` (async jobs with optional signed callbacks)
* [Speech Generation](https://docs.vichar.io/features/speech-generation) — `POST /v1/audio/speech`
* [Transcription](https://docs.vichar.io/features/transcription) — `POST /v1/audio/transcriptions`
# Using the AI SDK
URL: https://docs.vichar.io/developers/ai-sdk
Vichar ships a first-party provider for the [Vercel AI SDK](https://ai-sdk.dev): [`@llmgateway/ai-sdk-provider`](https://github.com/vicharai/api-ai-sdk-provider). One provider instance and one API key reach every model in the catalog.
## Install [#install]
```bash
pnpm add ai @llmgateway/ai-sdk-provider
```
Set your API key (create one from the [dashboard](https://app.vichar.io/dashboard)):
```bash
export LLM_GATEWAY_API_KEY=vichar_your_key_here
```
## Generate text [#generate-text]
```typescript
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { generateText } from "ai";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { text } = await generateText({
model: llmgateway("openai/gpt-4o"),
prompt: "Hello!",
});
```
Switching models is a one-line change — the same provider serves every model:
```typescript
const { text } = await generateText({
model: llmgateway("anthropic/claude-3-5-sonnet-20241022"),
prompt: "Hello!",
});
```
## Model ID formats [#model-id-formats]
Vichar supports two model ID formats:
* **Canonical model IDs** (`gpt-4o`) — smart routing picks the best provider based on uptime, throughput, price, and latency
* **Provider-prefixed IDs** (`openai/gpt-4o`) — routes to a specific provider with automatic failover if uptime drops below 90%
See the [routing documentation](https://docs.vichar.io/features/routing) for details and the [models page](https://app.vichar.io/dashboard) for the full catalog.
## Stream responses [#stream-responses]
```typescript
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { streamText } from "ai";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { textStream } = await streamText({
model: llmgateway("anthropic/claude-3-5-sonnet-20241022"),
prompt: "Write a poem about coding",
});
for await (const text of textStream) {
process.stdout.write(text);
}
```
## Next.js route handler [#nextjs-route-handler]
```typescript
// app/api/chat/route.ts
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { streamText } from "ai";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
export async function POST(req: Request) {
const { messages } = await req.json();
const result = await streamText({
model: llmgateway("openai/gpt-4o"),
messages,
});
return result.toDataStreamResponse();
}
```
## Tool calling [#tool-calling]
```typescript
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { generateText, tool } from "ai";
import { z } from "zod";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { text, toolResults } = await generateText({
model: llmgateway("openai/gpt-4o"),
tools: {
weather: tool({
description: "Get the weather for a location",
parameters: z.object({
location: z.string(),
}),
execute: async ({ location }) => {
return { temperature: 72, condition: "sunny" };
},
}),
},
prompt: "What's the weather in San Francisco?",
});
```
## Without the provider package [#without-the-provider-package]
If you prefer not to add a dependency, point `@ai-sdk/openai` at the gateway with a custom base URL:
```typescript
import { createOpenAI } from "@ai-sdk/openai";
import { generateText } from "ai";
const llmgateway = createOpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { text } = await generateText({
model: llmgateway("openai/gpt-4o"),
prompt: "Hello!",
});
```
Every request made through the AI SDK shows up in your
[Activity](https://docs.vichar.io/learn/activity) and [Usage & Metrics](https://docs.vichar.io/learn/usage-metrics)
dashboards like any other gateway request — with per-request cost, tokens, and
latency.
## Generate video [#generate-video]
Version 4 of the provider (AI SDK 7, Node.js 22 or later, ESM) adds `llmgateway.video()` for the SDK's experimental video API. The provider submits a gateway video job, the SDK polls it, and the finished file is downloaded from the authenticated content endpoint.
```typescript
import { writeFile } from "node:fs/promises";
import { llmgateway } from "@llmgateway/ai-sdk-provider";
import { experimental_generateVideo as generateVideo } from "ai";
const { video } = await generateVideo({
model: llmgateway.video("seedance-2-0"),
prompt: "A cinematic aerial view of ocean waves at sunrise",
duration: 8,
resolution: "1280x720",
poll: { intervalMs: 5_000, timeoutMs: 600_000 },
});
await writeFile("video.mp4", video.uint8Array);
```
`duration` and a text prompt are required. `resolution` maps to the gateway's `size`, `duration` to `seconds`, and `generateAudio` to `audio`. Pass a first frame as `prompt.image`, first and last frames with `frameImages`, and reference inputs with `inputReferences`; any other gateway field, such as `callback_url`, goes through `providerOptions.llmgateway`. `experimental_startVideo` and `experimental_getVideoStatus` let another process pick up a running job. Supported sizes, durations, and inputs depend on the model — see [video generation](https://docs.vichar.io/features/video-generation) for the REST reference. Projects on AI SDK 6 should stay on `@llmgateway/ai-sdk-provider@3`, which has no video support.
## Next steps [#next-steps]
* [Image generation with the AI SDK](https://docs.vichar.io/developers/ai-sdk-images)
* [Migrate from Vercel AI Gateway](https://docs.vichar.io/migrations/vercel-ai-gateway)
* [Reasoning support](https://docs.vichar.io/features/reasoning) and [caching](https://docs.vichar.io/features/caching)
# Vichar CLI
URL: https://docs.vichar.io/developers/cli
The **Vichar CLI** (`@llmgateway/cli`) is a command-line utility for launching coding agents pre-configured with Vichar, scaffolding projects, discovering models, and managing your Vichar account — API keys, spending budgets, and usage analytics — straight from the terminal.
## Installation [#installation]
Run commands directly without installation:
```bash
npx @llmgateway/cli init
```
Install globally for faster access:
```bash
npm install -g @llmgateway/cli
```
Then run commands directly (`lg` works as a shorthand alias):
```bash
llmgateway init
lg init
```
## Quick Start [#quick-start]
### Initialize a Project [#initialize-a-project]
Create a new project from a template:
```bash
npx @llmgateway/cli init
```
Or specify the template and name directly:
```bash
npx @llmgateway/cli init --template image-generation --name my-ai-app
```
### Sign In [#sign-in]
Sign in with your Vichar account to unlock key management, budgets, and usage analytics:
```bash
npx @llmgateway/cli auth login --email you@example.com
```
Or store a gateway API key only (enough for making gateway requests):
```bash
npx @llmgateway/cli auth login --key
```
Credentials are stored in `~/.llmgateway/config.json`. The `LLMGATEWAY_API_KEY` environment variable takes precedence over a stored key.
### Start Development [#start-development]
Navigate to your project and start the development server:
```bash
cd my-ai-app
npx @llmgateway/cli dev
```
Or specify a custom port:
```bash
npx @llmgateway/cli dev --port 3000
```
## Launch Coding Agents [#launch-coding-agents]
### `launch` [#launch]
Start any supported coding agent pre-wired to Vichar: one API key, 200+ models, and every request tracked in your [dashboard](https://app.vichar.io/dashboard).
```bash
# Interactive agent picker
npx @llmgateway/cli launch
# Launch a specific agent (shortcuts work too: `llmgateway claude`)
npx @llmgateway/cli launch claude
npx @llmgateway/cli launch opencode
npx @llmgateway/cli launch codex
# Pick a model — launcher flags go before the agent name
npx @llmgateway/cli launch -m gpt-5.5 claude
# Everything after the agent name is passed to the agent itself
npx @llmgateway/cli launch claude --continue
# List all supported agents and see which are installed
npx @llmgateway/cli launch --list
# Inspect what would run without launching
npx @llmgateway/cli launch --dry-run codex
```
Every supported agent also works as a direct shortcut, e.g. `npx @llmgateway/cli claude`. The launcher configures each agent automatically — environment variables, config files, or the agent's own key-registration command, whichever that agent needs — without overwriting your existing setup. For OpenCode and Claude Code, launching also applies the same model-catalog setup as [`configure`](#configure) on every launch, so their model pickers stay fresh as new models ship.
The API key is resolved from `--key`, the `LLMGATEWAY_API_KEY` environment variable, or the key stored by `llmgateway auth login --key` — in that order. Before launching, the key is verified against the gateway; a stale key (e.g. one you rolled or deleted) is reported with its exact source and the launcher falls back to the next valid one, prompting you for a fresh key if none works. If an agent isn't installed, the launcher prints its official install command and exits.
See the [integration guides](https://app.vichar.io/dashboard) for per-agent
setup details, and run `npx @llmgateway/cli launch --list` for the up-to-date
list of supported agents.
### `configure` [#configure]
Put Vichar's coding-model catalog directly into an agent's own config, so its model picker lists gateway models without launching through the CLI:
```bash
# opencode: adds every coding model, pinned per upstream provider, to the picker
npx @llmgateway/cli configure opencode
# Claude Code: routes it through Vichar and fills /model from the gateway catalog
npx @llmgateway/cli configure claude
# ...for the current repo only (.claude/settings.local.json)
npx @llmgateway/cli configure claude --project
# Preview without writing
npx @llmgateway/cli configure opencode --dry-run
```
* **OpenCode** — merges `provider/model` entries (e.g. `anthropic/claude-sonnet-5`, `aws-bedrock/claude-sonnet-5`) into `provider.llmgateway.models` in `~/.config/opencode/opencode.json`, with display names, context limits, and per-provider pricing. They show up in the picker as `llmgateway//` and pin that upstream provider via the gateway's [provider-routing syntax](https://docs.vichar.io/features/routing#provider-specific-routing). OpenCode now ships both catalogs natively too — canonical IDs under **Vichar** (`llmgateway`) and pinned `provider/model` IDs under **Vichar (provider-pinned)** (`llmgateway-providers`) — so `configure` mainly covers models the built-in catalog has not picked up yet. Existing custom entries and the rest of the file are preserved, and a hand-written `opencode.jsonc` keeps working alongside it.
* **Claude Code** — sets `ANTHROPIC_BASE_URL`, `ANTHROPIC_AUTH_TOKEN`, and `CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1` in `~/.claude/settings.json` (requires Claude Code v2.1.129+). Claude Code then loads the gateway's `/v1/models` catalog into its `/model` picker (shown as "From gateway"). Claude Code only lists IDs starting with `claude`/`anthropic`; any other gateway model still works via `claude --model `.
`llmgateway launch opencode` and `llmgateway launch claude` apply the same setup automatically on every launch, keeping the catalog fresh as new models ship.
## Project Commands [#project-commands]
### `init` [#init]
Initialize a new project from a template.
```bash
npx @llmgateway/cli init [directory] [options]
```
**Options:**
* `-t, --template ` — Template to use (default: `image-generation`)
* `-n, --name ` — Project name
**Examples:**
```bash
# Interactive mode
npx @llmgateway/cli init
# With options
npx @llmgateway/cli init --template image-generation --name my-app
```
### `list` [#list]
Display available project templates, grouped by category. Alias: `ls`.
```bash
npx @llmgateway/cli list
```
**Options:**
* `--json` — Output in JSON format
### `models` [#models]
Browse and filter available AI models.
```bash
npx @llmgateway/cli models [options]
```
**Options:**
* `-c, --capability ` — Filter by capability (e.g., `image`, `text`)
* `-p, --provider ` — Filter by provider (e.g., `openai`, `anthropic`)
* `-s, --search ` — Search models by name
* `--json` — Output in JSON format
**Examples:**
```bash
# List all models
npx @llmgateway/cli models
# Filter by provider
npx @llmgateway/cli models --provider openai
# Search models
npx @llmgateway/cli models --search gpt
```
### `add` [#add]
Add tools or API routes to an existing project.
```bash
npx @llmgateway/cli add [type] [name]
```
Runs interactively when `type` (`tool` or `route`) and `name` are omitted.
**Tools available:**
* `weather` — Weather lookup functionality
* `search` — Web search capability
* `calculator` — Mathematical operations
**API routes available:**
* `generate` — Text generation endpoint
* `chat` — Chat completion endpoint with streaming
### `dev` [#dev]
Start the local development server using your project's package manager.
```bash
npx @llmgateway/cli dev [options]
```
**Options:**
* `-p, --port ` — Port to run on
### `upgrade` [#upgrade]
Update Vichar dependencies (`@llmgateway/ai-sdk-provider`, `@llmgateway/models`, `@llmgateway/cli`) in your project.
```bash
npx @llmgateway/cli upgrade [options]
```
**Options:**
* `--check` — Check for updates without installing
### `docs` [#docs]
Open the documentation in your browser.
```bash
npx @llmgateway/cli docs [topic]
```
**Topics:** `models`, `api`, `sdk`, `quickstart` — omit to open the docs home and see all topics.
## Account Commands [#account-commands]
The commands below require a dashboard session — sign in first with
`llmgateway auth login --email`. A gateway API key alone is not enough for
account management.
### `auth` [#auth]
Manage authentication (dashboard session and gateway API key).
For browser-based device sign-in, the approval page is titled **Authorize your device**. Check that its code matches the one on your device before choosing **Authorize device**. Approval creates a separate account session; signing the device out leaves the browser signed in. Choose **Deny** for a request you did not start.
```bash
# Sign in with email & password (full access), or paste an API key
npx @llmgateway/cli auth login
npx @llmgateway/cli auth login --email you@example.com
npx @llmgateway/cli auth login --key
# Check authentication status (session + API key)
npx @llmgateway/cli auth status
# Show the signed-in user
npx @llmgateway/cli auth whoami
# Remove stored session and API key
npx @llmgateway/cli auth logout
```
### `keys` [#keys]
Create and manage gateway API keys.
```bash
npx @llmgateway/cli keys
```
#### `keys create` [#keys-create]
Create a new API key, optionally with spending limits and an expiry.
```bash
npx @llmgateway/cli keys create --description "CI key" --limit 100 --expires 30d
```
**Options:**
* `-p, --project ` — Project the key belongs to
* `-d, --description ` — Key description
* `-l, --limit ` — Total spending limit in USD (e.g. `100` or `49.99`)
* `--period-limit ` — Spending limit per rolling period in USD
* `--period ` — Rolling period for `--period-limit` (`12h`, `1d`, `2w`, `1mo`; default `1mo`)
* `-e, --expires ` — TTL as a duration (`30d`, `12h`) or an ISO date
* `--json` — Output in JSON format
The token is only displayed once at creation time — save it immediately.
#### `keys list` [#keys-list]
List API keys with spend, budget, and expiry. Alias: `keys ls`.
**Options:**
* `-p, --project ` — Filter by project
* `--all` — Show all keys in the org (admin/owner only)
* `--json` — Output in JSON format
#### `keys update ` [#keys-update-id]
Activate or deactivate an API key.
**Options:**
* `--activate` — Set the key to active
* `--deactivate` — Set the key to inactive
* `-e, --expires ` — New expiry as a duration (`30d`) or ISO date (needed to reactivate expired keys)
#### `keys limit ` [#keys-limit-id]
Set spending limits on an API key (same as `budget set`).
**Options:**
* `-l, --limit ` — Total spending limit in USD
* `--period-limit ` — Spending limit per rolling period in USD
* `--period ` — Rolling period (`12h`, `1d`, `2w`, `1mo`; default `1mo`)
* `--clear` — Remove all spending limits
#### `keys roll ` [#keys-roll-id]
Regenerate the token for an API key. The old token becomes invalid immediately.
**Options:**
* `-y, --yes` — Skip confirmation
#### `keys delete ` [#keys-delete-id]
Delete an API key. Alias: `keys rm`.
**Options:**
* `-y, --yes` — Skip confirmation
### `budget` [#budget]
Manage API key spending limits.
```bash
# Set a total and/or rolling-period budget
npx @llmgateway/cli budget set --limit 100 --period-limit 25 --period 1w
# Remove all spending limits
npx @llmgateway/cli budget set --clear
# Show budget and current spend
npx @llmgateway/cli budget get
```
**`budget set` options:** `-l, --limit `, `--period-limit `, `--period `, `--clear`
**`budget get` options:** `-p, --project `, `--json`
### `usage` [#usage]
View usage and cost analytics.
```bash
npx @llmgateway/cli usage [options]
```
**Options:**
* `-o, --org ` — Aggregate usage across an organization
* `-p, --project ` — Filter by project
* `-k, --api-key ` — Filter by API key
* `--by ` — Break down by `model` or `key`
* `-r, --range ` — Time range: `1h`, `4h`, `24h`, `7d`, `30d`, `365d` (default `7d`)
* `--days ` — Look back N days instead of `--range`
* `--from ` / `--to ` — Custom date range (`YYYY-MM-DD`)
* `--json` — Output in JSON format
**Examples:**
```bash
# Last 7 days for the default project
npx @llmgateway/cli usage
# Cost per model over the last 30 days
npx @llmgateway/cli usage --by model --range 30d
# Whole-org aggregate
npx @llmgateway/cli usage --org
```
#### `usage sources` [#usage-sources]
Break down usage by session/agent source to see which agents or sessions are spending.
```bash
npx @llmgateway/cli usage sources [options]
```
**Options:** `-p, --project `, `-r, --range ` (`7d`, `30d`), `--from `, `--to `, `--json`
### `orgs` [#orgs]
List your organizations with plan and credit balance. Alias: `orgs ls`.
```bash
npx @llmgateway/cli orgs list [--json]
```
### `projects` [#projects]
Manage projects and the CLI's default project.
```bash
# List projects (optionally filtered by org)
npx @llmgateway/cli projects list [--org ] [--json]
# Set the default project used by keys/budget/usage commands
npx @llmgateway/cli projects use
```
### `credits` [#credits]
Show organization credit balances.
```bash
npx @llmgateway/cli credits [--org ] [--json]
```
## Available Templates [#available-templates]
### Web Applications [#web-applications]
* **`image-generation`** — Full-stack AI image generation app (Next.js 16, React 19). Multi-provider support with a unified API.
* **`ai-chatbot`** — AI chatbot with streaming responses.
* **`og-image-generator`** — AI-powered OG image generator.
* **`feedback-dashboard`** — Customer feedback sentiment dashboard.
* **`writing-assistant`** — AI writing assistant with text actions.
* **`qa-agent`** — AI-powered QA testing agent with browser automation, real-time action timeline, and live browser preview.
* **`showcase`** — Public, filterable gallery of apps built with Vichar templates. Static and deployable, with a "Submit your app" flow.
### Bots [#bots]
* **`slack-qa-bot`** — Slack bot that streams AI answers and keeps thread context.
### CLI Agents [#cli-agents]
* **`weather-agent`** — Answers weather queries using tool calling.
* **`lead-agent`** — Researches people and posts results to Discord.
* **`changelog-generator-agent`** — Generates changelogs from git history.
* **`email-drafter-agent`** — Drafts polished emails from rough notes.
* **`sentiment-analyzer-agent`** — Analyzes text sentiment.
* **`data-extractor-agent`** — Extracts structured entities from text.
```bash
npx @llmgateway/cli init --template qa-agent
```
## Configuration [#configuration]
The CLI stores configuration in `~/.llmgateway/config.json`:
```json
{
"apiKey": "vichar_...",
"defaultTemplate": "image-generation",
"sessionEmail": "you@example.com",
"defaultOrgId": "org_...",
"defaultProjectId": "proj_..."
}
```
Signing in with `auth login --email` also stores a dashboard session used by the account commands (`keys`, `budget`, `usage`, `orgs`, `projects`, `credits`).
### Environment Variables [#environment-variables]
* `LLMGATEWAY_API_KEY` — Gateway API key; takes precedence over the config file:
```bash
export LLMGATEWAY_API_KEY="vichar_..."
```
* `LLMGATEWAY_API_URL` — Override the management API base URL (defaults to `https://api.vichar.io`), useful for self-hosted deployments.
## More Resources [#more-resources]
* [GitHub Repository](https://github.com/vicharai/api-templates) — Source code and issues
Need help or want to request a feature? Open an issue on
[GitHub](https://github.com/vicharai/api-templates/issues).
# Vichar Developer Resources
URL: https://docs.vichar.io/developers
This section is for developers building applications on top of Vichar — with our command-line tool, the MCP server, the [Vercel AI SDK](https://ai-sdk.dev) via our first-party provider package [`@llmgateway/ai-sdk-provider`](https://github.com/vicharai/api-ai-sdk-provider), and [TanStack AI](https://tanstack.com/ai) via the first-party [`@tanstack/ai-llmgateway`](https://www.npmjs.com/package/@tanstack/ai-llmgateway) adapter.
## API entry points [#api-entry-points]
* [OpenAPI specification](https://api.vichar.io/openapi.json) — typed request, response, and error schemas
* [Authentication](https://docs.vichar.io/features/api-keys) — API keys and access control
* [Developer dashboard](https://app.vichar.io/dashboard) — projects, keys, usage, and budgets
* [API versioning and deprecation policy](https://docs.vichar.io/resources/api-versioning) — compatibility and retirement notices
## Guides [#guides]
* [**Vichar CLI**](https://docs.vichar.io/developers/cli) — Launch coding agents, scaffold projects from templates, generate agent configs, and manage keys, budgets, and usage from the terminal
* [**Model Context Protocol (MCP)**](https://docs.vichar.io/developers/mcp) — Use Vichar as an MCP server from Claude Code, Cursor, and other MCP clients
* [**Using the AI SDK**](https://docs.vichar.io/developers/ai-sdk) — Install the provider, generate and stream text, call tools, and wire up Next.js routes
* [**Image Generation with the AI SDK**](https://docs.vichar.io/developers/ai-sdk-images) — Generate images with `generateImage` and stream image output through chat
* [**Using TanStack AI**](https://docs.vichar.io/developers/tanstack-ai) — Stream chat with `useChat`, call tools, and surface reasoning through the first-party `@tanstack/ai-llmgateway` adapter
## Why the AI SDK [#why-the-ai-sdk]
The AI SDK gives you one TypeScript interface for text generation, streaming, tool calling, and image generation. Combined with Vichar, a single provider instance and one API key reach every model in the catalog — see the [models page](https://app.vichar.io/dashboard) for what's available.
## Other ways to integrate [#other-ways-to-integrate]
If you're not using the AI SDK:
* Use any OpenAI-compatible SDK against `https://api.vichar.io/v1` — see the [Quickstart](https://docs.vichar.io/quick-start)
* Use the Anthropic SDK against the [Anthropic-compatible endpoint](https://docs.vichar.io/features/anthropic-endpoint)
* Call the REST API directly — see the API reference in the sidebar
## Brand assets [#brand-assets]
For integration listings, presentations, and partner pages, download the
official \[Vichar brand assets] in SVG or transparent PNG.
The brand guide covers the horizontal logo, standalone symbol, clear space,
minimum sizes, background colors, and typography.
# Vichar MCP Server
URL: https://docs.vichar.io/developers/mcp
Connect your AI assistant to Vichar to inspect your usage and costs, discover your most-used models, providers and coding apps, and generate text or images. The same API key connects all of these tools.
## Video walkthrough [#video-walkthrough]
## Connection and discovery [#connection-and-discovery]
Connect with Streamable HTTP at `https://api.vichar.io/mcp`. Send an API key in `Authorization: Bearer `. A GET without an SSE Accept header returns public server information;.
For protocol requests, POST a single JSON-RPC message with `Content-Type: application/json` and `Accept: application/json, text/event-stream`. Initialize first, then send the negotiated `MCP-Protocol-Version` header on subsequent requests. The transport is stateless: requests return JSON, accepted notifications return an empty 202, and standalone SSE subscriptions and session deletion return 405. Clients using the original HTTP+SSE bridge remain supported.
[Protected resource metadata](https://api.vichar.io/.well-known/oauth-protected-resource/mcp) publishes authentication discovery. Unauthorized protocol requests include a `WWW-Authenticate` challenge pointing to it.
## What is MCP? [#what-is-mcp]
The Model Context Protocol (MCP) is an open standard that allows AI assistants to connect with external tools and data sources. Vichar's MCP server exposes tools for:
* **Account and usage analytics** - Check spending limits, request/token totals, costs, trends, and provider/model/app rankings
* **Chat completions** - Send messages to any supported LLM
* **Image generation** - Generate images using models like Qwen Image
* **Nano Banana image generation** - Generate images with Gemini 3 Pro Image and optionally save to disk
* **Model discovery** - List available models with capabilities and pricing
## Available Tools [#available-tools]
### `get-account` [#get-account]
Inspect the connected user, organization, project, role, analytics scope, and API key spending limits. Owners and admins also receive the organization's current credit balance. No parameters are required, and credentials are never returned.
### `get-usage` [#get-usage]
Get request/token totals, errors, cache hits, costs, a time series, and your most-used provider, model, and coding agent/app **by request count**.
| Parameter | Description |
| ------------- | ---------------------------------------------------------------------------------- |
| `from` | Optional first UTC date, `YYYY-MM-DD`, inclusive. Defaults to 29 days before `to`. |
| `to` | Optional last UTC date, inclusive. Defaults to today. Maximum range: 366 days. |
| `granularity` | `day` (default) or `hour`. Hourly reports allow at most 31 days. |
```json
{
"from": "2026-08-01",
"to": "2026-08-31",
"granularity": "day"
}
```
The response includes `scope`, the resolved dates, `updatedAt`, `totals`, `series`, `mostUsedProvider`, `mostUsedModel`, `mostUsedApp`, and `appUsageCoverage`. Only time buckets with activity appear in `series`. Empty periods return zero totals, an empty series, and null rankings.
### `get-usage-breakdown` [#get-usage-breakdown]
Rank providers, models, coding apps, or API keys by requests, inference cost, or tokens.
| Parameter | Description |
| ------------ | ----------------------------------------------------------------------- |
| `group_by` | Required: `provider`, `model`, `app`, or `api_key`. |
| `sort_by` | `requests` (default), `cost`, or `tokens`, descending. Ties use the ID. |
| `from`, `to` | Same inclusive UTC dates as `get-usage`. |
| `limit` | Results per page: 1–100, default 10. |
| `offset` | Results to skip: 0–10000, default 0. |
```json
{
"group_by": "app",
"sort_by": "cost",
"limit": 10
}
```
The response includes each row's ID, display name, requests, tokens and costs, plus `pagination.hasMore` and `coverage`. Increase `offset` by `limit` to get the next page. Known app aliases are combined before ranking. `unknown` identifies requests with no recorded source; custom app names remain as recorded.
### Analytics scope and cost fields [#analytics-scope-and-cost-fields]
* Owners and admins see the **connected project's** usage across its keys. Developers see only the keys they created in that project, including inactive-key history.
* A project API key does not grant access to another project or organization. Tools do not accept scope overrides. Use a key for the project you want to inspect.
* Use an active user API key. Customer credentials and expired or revoked keys cannot read account analytics. Project access is checked on every analytics request.
* Analytics tools are read-only, incur no model charges, and remain available when a key or member reaches a spending limit. Generation tools still enforce those limits.
* `costUsd` is inference usage cost. `creditsCostUsd` and `byokCostUsd` separate gateway credits from provider costs paid with your own provider keys. `dataStorageCostUsd` is separate. These are usage statistics, not invoice totals or exact changes in credit balance.
* Statistics come from hourly aggregates, survive request-retention cleanup, and may lag recent requests. `updatedAt` reports the last summary aggregation in the requested period.
* App attribution uses the request's recorded source, including recognized coding clients and `x-source` values. MCP generation calls preserve client attribution headers. Configure `x-source` on your MCP connection if your client does not identify itself. Attribution does not identify which person used a shared key.
* Historical per-key app statistics start when per-key source aggregation is enabled. `appUsageCoverage` / `coverage` compare recorded breakdown requests with total requests; `complete: false` means rankings cover only part of the period. This differs from an `unknown` source, which is a recorded request without app attribution.
All three tools return JSON in both `structuredContent` and a text content block for older clients. An unavailable backend produces a tool error, never a fabricated zero-usage report.
### `chat` [#chat]
Send a message to any LLM and get a response.
**Parameters:**
* `model` (string) - A model ID from `list-models` or the [live catalog](https://app.vichar.io/dashboard)
* `messages` (array) - Array of messages with `role` and `content`
* `temperature` (number, optional) - Sampling temperature (0-2)
* `max_tokens` (number, optional) - Maximum tokens to generate
**Example:**
```json
{
"model": "MODEL_ID",
"messages": [{ "role": "user", "content": "Explain quantum computing" }],
"temperature": 0.7
}
```
### `generate-image` [#generate-image]
Generate images from text prompts using AI image models.
**Parameters:**
* `prompt` (string) - Text description of the image to generate
* `model` (string, optional) - Image model (default: `"qwen-image-3.0"`)
* `size` (string, optional) - Image size (default: `"1024x1024"`)
* `n` (number, optional) - Number of images (1-4, default: 1)
**Example:**
```json
{
"prompt": "A serene mountain landscape at sunset",
"model": "qwen-image-3.0-pro",
"size": "1024x1024"
}
```
### `generate-nano-banana` [#generate-nano-banana]
Generate an image using Gemini 3 Pro Image ("Nano Banana Pro"). Returns an inline image preview, and optionally saves the image to disk when the server is configured with an upload directory.
**Parameters:**
* `prompt` (string) - Text description of the image to generate
* `filename` (string, optional) - Filename for the saved image, no path separators allowed (default: `nano-banana-{timestamp}.png`)
* `aspect_ratio` (string, optional) - Aspect ratio: `"1:1"`, `"16:9"`, `"4:3"`, or `"5:4"`
**Example:**
```json
{
"prompt": "A pixel-art cat sitting on a rainbow",
"filename": "hero-image.png",
"aspect_ratio": "16:9"
}
```
**Saving images to disk** requires the `UPLOAD_DIR` environment variable to be
set on the MCP server. When set, images are saved to that directory. Without
it, images are returned inline only — no files are written to disk. See
[Enabling local image saving](#enabling-local-image-saving) for setup
instructions.
### `list-models` [#list-models]
List available LLM models with capabilities and pricing.
**Parameters:**
* `include_deactivated` (boolean, optional) - Include deactivated models
* `exclude_deprecated` (boolean, optional) - Exclude deprecated models
* `limit` (number, optional) - Maximum models to return (default: 20)
* `family` (string, optional) - Filter by model family
### `list-image-models` [#list-image-models]
List all available image generation models.
Use the tool for current model IDs, capabilities, and pricing, or browse the [live models page](https://app.vichar.io/dashboard).
## Setup [#setup]
### Get Your API Key [#get-your-api-key]
1. Log in to your [Vichar dashboard](https://app.vichar.io/dashboard)
2. Navigate to **API Keys** section
3. Create a new API key and copy it
### Configure Claude Code [#configure-claude-code]
Run the following command in your terminal:
```bash
claude mcp add --transport http --scope user llmgateway https://api.vichar.io/mcp \
--header "Authorization: Bearer your-api-key-here"
```
**Alternative: Manual configuration**
You can also add the MCP server manually by editing `~/.claude.json` (user scope) or `.mcp.json` in your project root (project scope):
```json
{
"mcpServers": {
"llmgateway": {
"url": "https://api.vichar.io/mcp",
"headers": {
"Authorization": "Bearer your-api-key-here"
}
}
}
}
```
Restart Claude Code after manual configuration changes.
### Test the Integration [#test-the-integration]
Try using the tools in Claude Code:
* "Show my usage and costs for the last 30 days"
* "Generate an image of a futuristic city using the generate-image tool"
* "Use generate-nano-banana to create a hero image for my landing page"
* "Which model, provider, and coding app do I use most?"
### Get Your API Key [#get-your-api-key-1]
1. Log in to your [Vichar dashboard](https://app.vichar.io/dashboard)
2. Navigate to **API Keys** section
3. Create a new API key and copy it
4. Set it as an environment variable: `export LLM_GATEWAY_API_KEY="your-api-key-here"`
### Configure Codex [#configure-codex]
Run the following command in your terminal:
```bash
codex mcp add llmgateway --url https://api.vichar.io/mcp \
--bearer-token-env-var LLM_GATEWAY_API_KEY
```
**Alternative: Manual configuration**
You can also add the MCP server manually by editing `~/.codex/config.toml`:
```toml
[mcp_servers.llmgateway]
url = "https://api.vichar.io/mcp"
bearer_token_env_var = "LLM_GATEWAY_API_KEY"
```
### Test the Integration [#test-the-integration-1]
Run `/mcp` in the Codex TUI to confirm the `llmgateway` server is connected. Try:
* "Show my usage and costs for the last 30 days"
* "Generate an image of a futuristic city using the generate-image tool"
* "Use generate-nano-banana to create a hero image for my landing page"
* "Which model, provider, and coding app do I use most?"
### Get Your API Key [#get-your-api-key-2]
1. Log in to your [Vichar dashboard](https://app.vichar.io/dashboard)
2. Navigate to **API Keys** section
3. Create a new API key and copy it
### Configure Cursor [#configure-cursor]
Add the following to your Cursor MCP configuration file (`~/.cursor/mcp.json`):
```json
{
"mcpServers": {
"llmgateway": {
"url": "https://api.vichar.io/mcp",
"headers": {
"Authorization": "Bearer your-api-key-here"
}
}
}
}
```
Or open the Command Palette (`Cmd/Ctrl + Shift + P`), search for **"Cursor Settings"**, then go to **Tools & Integrations** > **Add Custom MCP** and paste the configuration above.
Cursor v0.48.0+ is required for Streamable HTTP MCP support.
### Test the Integration [#test-the-integration-2]
Open a chat in **Agent Mode**, click the **Select Tools** icon, and verify the Vichar tools appear. Try:
* "Show my usage and costs for the last 30 days"
* "Generate an image of a futuristic city using the generate-image tool"
* "Use generate-nano-banana to create a hero image for my landing page"
* "Which model, provider, and coding app do I use most?"
Vichar's MCP server supports the standard HTTP Streamable transport. Configure your client with:
* **Endpoint:** `https://api.vichar.io/mcp`
* **Authentication:** Bearer token via `Authorization` header or `x-api-key` header
* **Protocol Version:** 2024-11-05
**Direct HTTP Example:**
```bash
curl -X POST https://api.vichar.io/mcp \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-api-key" \
-d '{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list"
}'
```
**Server-Sent Events (SSE):**
For real-time updates, connect with `Accept: text/event-stream`:
```bash
curl -N https://api.vichar.io/mcp \
-H "Accept: text/event-stream" \
-H "Authorization: Bearer your-api-key"
```
## Verify with MCP Inspector [#verify-with-mcp-inspector]
For a standalone connection check, start the official [MCP Inspector](https://modelcontextprotocol.io/docs/tools/inspector):
```bash
pnpm dlx @modelcontextprotocol/inspector
```
In Inspector, add a server manually:
1. Choose **Streamable HTTP** and enter `https://api.vichar.io/mcp`.
2. In the server settings, add a custom `Authorization` header with `Bearer YOUR_API_KEY`.
3. Connect, open **Tools**, and select **chat**.
4. Enter an accessible text model ID from the [live catalogue](https://app.vichar.io/dashboard), a `messages` array, and an optional `max_tokens` limit.
5. Click **Execute Tool** and inspect the returned text and token usage.
For example, the `messages` field accepts:
```json
[
{
"role": "user",
"content": "In one sentence, explain what an API gateway does."
}
]
```
Keep `stream` disabled: this MCP tool returns a complete response. Model generation is billed to the workspace that issued the key. Use **list-models** to discover available model metadata without generating a response.
Use HTTPS for the hosted endpoint. A self-hosted gateway also needs an HTTPS `MCP_GATEWAY_URL` for authenticated generation calls.
## Use Cases [#use-cases]
### Usage and Spending [#usage-and-spending]
```text
What did I spend this month, and which coding app accounts for the most cost?
Show my most-used model and provider, then compare daily usage with last month.
```
Use `get-account` to confirm the scope, `get-usage` for each period, and `get-usage-breakdown` with `group_by: "app"` and `sort_by: "cost"` for the app ranking.
### Multi-Model Access in Claude Code [#multi-model-access-in-claude-code]
Use Claude Code to interact with models it doesn't natively support:
```
List available models, then use the chat tool with a suitable model to review this code.
```
### Image Generation [#image-generation]
Generate images directly from your AI assistant:
```
Use generate-image to create a logo for my new startup.
It should be minimalist, blue and white, representing AI and cloud computing.
```
### Nano Banana (Gemini Image Generation) [#nano-banana-gemini-image-generation]
Generate images with Gemini 3 Pro for use in your project:
```
Use generate-nano-banana to create a hero image for my landing page with a 16:9 aspect ratio.
```
### Cost-Effective Model Selection [#cost-effective-model-selection]
Query available models to find the best option for your task:
```
List models and their pricing, then choose a suitable low-cost model for this task.
```
## Authentication [#authentication]
The MCP server supports two authentication methods:
1. **Bearer Token** - `Authorization: Bearer your-api-key`
2. **API Key Header** - `x-api-key: your-api-key`
Use the same project API key you use for inference. Analytics access follows the [scope rules above](#analytics-scope-and-cost-fields).
## OAuth Support [#oauth-support]
For applications that prefer OAuth authentication, Vichar's MCP server implements OAuth 2.0:
* **Authorization Endpoint:** `/oauth/authorize`
* **Token Endpoint:** `/oauth/token`
* **Registration Endpoint:** `/oauth/register`
* **Supported Flows:** Authorization Code, Client Credentials
## Enabling Local Image Saving [#enabling-local-image-saving]
By default, `generate-nano-banana` returns images inline without writing to disk. To enable saving generated images to the server filesystem, the `UPLOAD_DIR` environment variable must be set on the **gateway host** at startup. This is a server-side setting — it cannot be configured from the client.
This is only possible for **self-hosted** MCP deployments. Configure `UPLOAD_DIR` using your deployment method:
* **Docker:** Pass `-e UPLOAD_DIR=/data/images` or add it to your `docker-compose.yml` environment section.
* **systemd:** Add `Environment=UPLOAD_DIR=/data/images` to your service unit file.
* **.env file:** Add `UPLOAD_DIR=/data/images` to the `.env` file loaded by your gateway process.
The shared hosted endpoint (`api.vichar.io`) does not support configuring
`UPLOAD_DIR`. On the hosted service, images are always returned inline — no
files are written to disk. To enable server-side image saving, you must
self-host the MCP server and set `UPLOAD_DIR` at startup.
## Troubleshooting [#troubleshooting]
### Connection Errors [#connection-errors]
If you're having trouble connecting:
1. Verify your API key is valid
2. Check the endpoint URL is correct: `https://api.vichar.io/mcp`
3. Ensure your firewall allows outbound HTTPS connections
### Tool Not Found [#tool-not-found]
If tools aren't appearing:
1. Restart your MCP client
2. Check the configuration syntax
3. Verify the MCP server is responding: `GET https://api.vichar.io/mcp`
### Rate Limiting [#rate-limiting]
The MCP server respects your account's rate limits. If you're hitting limits:
1. Check your usage in the dashboard
2. Consider upgrading your plan
3. Implement request queuing in your application
Need help? Join our [support](mailto:contact@vichar.io) for support.
## Benefits [#benefits]
* **Unified Access** - Discover the [live model and provider catalog](https://app.vichar.io/dashboard) through one interface
* **Cost Tracking** - Ask your assistant about usage, spending, and your most-used models, providers, and apps
* **Caching** - Automatic response caching reduces costs and latency
* **Fallback** - Automatic provider failover ensures reliability
* **Image Generation** - Generate images directly from your AI assistant
# Using TanStack AI
URL: https://docs.vichar.io/developers/tanstack-ai
[TanStack AI](https://tanstack.com/ai) ships a first-party Vichar adapter: [`@tanstack/ai-llmgateway`](https://www.npmjs.com/package/@tanstack/ai-llmgateway), maintained in the TanStack AI repository alongside the OpenAI and Anthropic adapters. One adapter and one API key reach every model in the catalog.
## Install [#install]
```bash
pnpm add @tanstack/ai @tanstack/ai-react @tanstack/ai-llmgateway
```
`@tanstack/ai-react` is the React client; TanStack AI also ships Vue, Svelte, Angular, and Preact packages that work with the same adapter.
Set your API key (create one from the [dashboard](https://app.vichar.io/dashboard)):
```bash
export LLM_GATEWAY_API_KEY=vichar_your_key_here
```
## Stream chat from a server route [#stream-chat-from-a-server-route]
`llmGatewayText(model)` creates the adapter and reads the key from `LLM_GATEWAY_API_KEY`:
```typescript
// app/api/chat/route.ts
import { chat, toServerSentEventsResponse } from "@tanstack/ai";
import { llmGatewayText } from "@tanstack/ai-llmgateway";
export async function POST(request: Request) {
const { messages } = await request.json();
const stream = chat({
adapter: llmGatewayText("gpt-5.6-terra"),
messages,
});
return toServerSentEventsResponse(stream);
}
```
Switching models is a one-line change — the same adapter serves every model:
```typescript
adapter: llmGatewayText("claude-sonnet-5"),
```
## Model ID formats [#model-id-formats]
Vichar supports two model ID formats:
* **Canonical model IDs** (`gpt-5.6-terra`) — smart routing picks the best provider based on uptime, throughput, price, and latency
* **Provider-prefixed IDs** (`moonshot/kimi-k3`) — routes to a specific provider with automatic failover if uptime drops below 90%
A curated set of flagship models carries typed metadata (`LLMGATEWAY_CHAT_MODELS`) with editor autocomplete for input modalities and options; any other ID from the [models page](https://app.vichar.io/dashboard) still works. See the [routing documentation](https://docs.vichar.io/features/routing) for details.
## Connect the React client [#connect-the-react-client]
`useChat` consumes the AG-UI event stream from the route above — no client-side API key, no per-provider wiring:
```tsx
// components/chat.tsx
"use client";
import { fetchServerSentEvents, useChat } from "@tanstack/ai-react";
import { useState } from "react";
export function Chat() {
const [input, setInput] = useState("");
const { messages, sendMessage, isLoading } = useChat({
connection: fetchServerSentEvents("/api/chat"),
});
return (
);
}
```
## Tool calling [#tool-calling]
Define tools with `toolDefinition` and attach a server handler — TanStack AI runs the tool loop for you:
```typescript
import { chat, toServerSentEventsResponse, toolDefinition } from "@tanstack/ai";
import { llmGatewayText } from "@tanstack/ai-llmgateway";
import { z } from "zod";
const getWeather = toolDefinition({
name: "get_weather",
description: "Get the current weather for a location",
inputSchema: z.object({
location: z.string(),
}),
}).server(async ({ location }) => {
return { temperature: 72, condition: "sunny" };
});
export async function POST(request: Request) {
const { messages } = await request.json();
const stream = chat({
adapter: llmGatewayText("gpt-5.6-terra"),
messages,
tools: [getWeather],
});
return toServerSentEventsResponse(stream);
}
```
## Reasoning models [#reasoning-models]
Reasoning models stream their thinking as `reasoning_content` deltas, which the adapter surfaces as AG-UI `REASONING_*` events — they arrive in `useChat` as `thinking` parts. Control the depth with `reasoning_effort` in `modelOptions`:
```typescript
const stream = chat({
adapter: llmGatewayText("kimi-k3"),
messages,
modelOptions: {
temperature: 0.7,
reasoning_effort: "high",
},
});
```
`reasoning_effort` accepts the extended scale `none` / `minimal` / `low` / `medium` / `high` / `xhigh` / `max`; which tiers a model honors depends on the model and the provider it routes to. Parameters a routed provider doesn't support are stripped server-side, so `modelOptions` stay portable across models. See [reasoning support](https://docs.vichar.io/features/reasoning).
## Summarization [#summarization]
The adapter also covers TanStack AI's `summarize` surface:
```typescript
import { summarize } from "@tanstack/ai";
import { llmGatewaySummarize } from "@tanstack/ai-llmgateway";
const result = await summarize({
adapter: llmGatewaySummarize("gpt-5.4-mini"),
text: "Long article text...",
stream: false,
});
```
## Self-hosted deployments [#self-hosted-deployments]
`createLLMGatewayText` takes the key explicitly plus a `baseURL` for self-hosted gateways:
```typescript
import { createLLMGatewayText } from "@tanstack/ai-llmgateway";
const adapter = createLLMGatewayText(
"gpt-5.6-terra",
process.env.LLM_GATEWAY_API_KEY!,
{
baseURL: "https://gateway.internal.example.com/v1",
},
);
```
Every request made through TanStack AI shows up in your
[Activity](https://docs.vichar.io/learn/activity) and [Usage & Metrics](https://docs.vichar.io/learn/usage-metrics)
dashboards like any other gateway request — with per-request cost, tokens, and
latency.
## Next steps [#next-steps]
* [TanStack AI adapter reference](https://tanstack.com/ai/latest/docs/adapters/llmgateway)
* [Routing and fallback](https://docs.vichar.io/features/routing)
* [Reasoning support](https://docs.vichar.io/features/reasoning) and [caching](https://docs.vichar.io/features/caching)
# Anthropic API Compatibility
URL: https://docs.vichar.io/features/anthropic-endpoint
Vichar provides a native Anthropic-compatible endpoint at `/v1/messages` that allows you to use any model in our catalog while maintaining the familiar Anthropic API format
This is especially useful for applications designed for Claude that you want to extend to use other models.
Enjoy a 50% discount on our Anthropic models for a limited time.
## Overview [#overview]
The Anthropic endpoint transforms requests from Anthropic's message format to the OpenAI-compatible format used by Vichar, then transforms the responses back to Anthropic's format. This means you can:
* Use **any model** available in Vichar with Anthropic's API format
* Maintain existing code that uses Anthropic's SDK or API format
* Access models from OpenAI, Google, Cohere, and other providers through the Anthropic interface
* Leverage Vichar's routing, caching, and cost optimization features
## Basic Usage [#basic-usage]
## Configuration for Claude Code [#configuration-for-claude-code]
This endpoint is perfect for configuring Claude Code to use any model available in Vichar:
```bash
export ANTHROPIC_BASE_URL=https://api.vichar.io
export ANTHROPIC_AUTH_TOKEN=vichar_your_api_key_here
# optional: specify a model, otherwise it uses the default Claude model
export ANTHROPIC_MODEL=gpt-5 # or any model from our catalog
# now run claude!
claude
```
Environment variables are read once at startup. The `/model` picker lists
Claude models only, so non-Claude models are selected with `ANTHROPIC_MODEL`
or `--model`. See the [Claude Code guide](https://docs.vichar.io/guides/claude-code) for the
settings-file options and gateway model discovery.
### Choosing Models [#choosing-models]
You can use any model from the [models page](https://app.vichar.io/dashboard). Popular options for Claude Code include:
```bash
# Use OpenAI's latest model
export ANTHROPIC_MODEL=gpt-5
# Use a cost-effective alternative
export ANTHROPIC_MODEL=gpt-5-mini
# Use Google's Gemini
export ANTHROPIC_MODEL=gemini-3.1-pro-preview
# Use Anthropic's actual Claude models
export ANTHROPIC_MODEL=claude-3-5-sonnet-20241022
```
## Environment Variables [#environment-variables]
When configuring Claude Code or other Anthropic-compatible applications, you can use these environment variables:
### ANTHROPIC\_MODEL [#anthropic_model]
Specifies the main model to use for primary requests.
* **Default**: `claude-sonnet-4-20250514`
* **Example**: `export ANTHROPIC_MODEL=gpt-5`
### ANTHROPIC\_SMALL\_FAST\_MODEL [#anthropic_small_fast_model]
Specifies a smaller, faster model used for background functionality and internal operations.
* **Default**: `claude-3-5-haiku-20241022`
* **Example**: `export ANTHROPIC_SMALL_FAST_MODEL=gpt-5-nano`
```bash
# Example configuration
export ANTHROPIC_BASE_URL=https://api.vichar.io
export ANTHROPIC_AUTH_TOKEN=vichar_your_api_key_here
export ANTHROPIC_MODEL=gpt-5
export ANTHROPIC_SMALL_FAST_MODEL=gpt-5-nano
```
## Advanced Features [#advanced-features]
### Making a manual request [#making-a-manual-request]
```bash
curl -X POST "https://api.vichar.io/v1/messages" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
],
"max_tokens": 100
}'
```
### Response Format [#response-format]
The endpoint returns responses in Anthropic's message format:
```json
{
"id": "msg_abc123",
"type": "message",
"role": "assistant",
"model": "gpt-5",
"content": [
{
"type": "text",
"text": "Hello! I'm doing well, thank you for asking. How can I help you today?"
}
],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 13,
"output_tokens": 20
}
}
```
### Request Format [#request-format]
`/v1/messages` expects Anthropic Messages requests and always answers in Anthropic's format. Because the two formats share `model` and `messages`, an OpenAI Chat Completions body can reach this endpoint by accident — and since unknown parameters are ignored rather than rejected, the request succeeds and returns an Anthropic response body that OpenAI SDKs cannot read. If your client reports an empty completion here, check that it is pointed at `/v1/chat/completions`.
Unknown parameters are deliberately ignored rather than rejected, so a valid Anthropic request is never denied for carrying an extra field. As a consequence, OpenAI-only parameters (`response_format`, `stream_options`, `max_completion_tokens`, `n`, `stop`, `seed`, `frequency_penalty`, and similar) have no effect here — the model will not honour them. Use `/v1/chat/completions` if you need them.
A body that is *structurally* OpenAI is rejected by the schema, as it always has been — OpenAI-shaped `tools` (`{"type": "function", "function": {…}}`), OpenAI content parts such as `image_url`, or assistant turns with `content: null`. Those rejections now name the mismatch and point at the right endpoint instead of returning an opaque validation error:
```json
{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "This endpoint implements Anthropic's Messages API, and the request body uses OpenAI Chat Completions structures (tools[0].function) that Anthropic's format has no equivalent for. Send OpenAI-format requests to /v1/chat/completions instead, or convert the body to Anthropic's Messages format."
}
}
```
Rejected requests are recorded in your logs with a `client_error` finish reason and zero cost, so a malformed client is visible in the activity feed rather than failing silently.
### Prompt Caching [#prompt-caching]
For Claude models, `cache_control` markers on `system` and message content blocks are forwarded to the provider unchanged, including the optional `ttl` (`5m` or `1h`):
```json
{
"model": "claude-sonnet-4-6",
"max_tokens": 100,
"system": [
{
"type": "text",
"text": "",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [{ "role": "user", "content": "Hello!" }]
}
```
Cache usage comes back in Anthropic's native fields: `usage.cache_creation_input_tokens` (tokens written to the cache this request, billed at the write premium), `usage.cache_read_input_tokens` (tokens served from cache at the discounted rate), and `usage.cache_creation` (the per-TTL write breakdown).
Each Claude model has a minimum cacheable prompt length (it varies by model
and is exposed as `min_cacheable_tokens` on `/v1/models`). A `cache_control`
marker on a shorter prompt is accepted but silently not cached — both cache
usage fields stay `0`. See [Provider Cache
Control](https://docs.vichar.io/features/caching/provider-cache-control) for details.
### Web Search [#web-search]
Anthropic's server-side web search tool works on this endpoint. Pass it as usual and the response carries `server_tool_use` and `web_search_tool_result` blocks before the text that cites them, so Anthropic SDK clients surface sources:
```json
{
"model": "claude-haiku-4-5",
"max_tokens": 400,
"messages": [{ "role": "user", "content": "What shipped in Node 24?" }],
"tools": [{ "type": "web_search_20250305", "name": "web_search" }]
}
```
Replaying the assistant turn verbatim on the next request is supported: the `server_tool_use` and `web_search_tool_result` blocks are accepted and dropped, since the provider re-runs the search rather than reusing the previous results.
### Tool Search [#tool-search]
Anthropic's server-side [tool search](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool) works on this endpoint. Pass a `tool_search_tool_*` tool alongside your catalog and mark the tools that should load on demand with `defer_loading: true`:
```json
{
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "What is the weather in Paris?" }],
"tools": [
{
"type": "tool_search_tool_regex_20251119",
"name": "tool_search_tool_regex"
},
{
"name": "get_weather",
"description": "Get the weather at a specific location",
"input_schema": { "type": "object" },
"defer_loading": true
}
]
}
```
Deferred tools stay out of the rendered tools section, so adding one does not invalidate an existing prompt cache. The response carries the `server_tool_use` and `tool_search_tool_result` blocks; replay them verbatim on the next request and Anthropic keeps expanding the `tool_reference` entries they carry, so Claude reuses a discovered tool instead of searching again. `tool_reference` blocks returned from your own client-side search inside a `tool_result` are forwarded unchanged too.
Send every tool definition on every request, including the deferred ones — Anthropic needs them server-side to run the search. At least one tool must stay non-deferred (normally the tool search tool itself), and a tool cannot carry both `defer_loading: true` and `cache_control`.
**Where it works.** Tool search reaches the provider on the Anthropic API and on Anthropic models served through Google Cloud, and it needs a Claude 4.5-generation model or newer — older Claude models reject it upstream. On every other provider, including Anthropic models on AWS Bedrock, the tool search tool and `defer_loading` are dropped and all tools are sent eagerly. The request still succeeds, it just loses the cache and token savings, so pin the provider (`anthropic/claude-sonnet-4-6`) when those savings matter.
Bedrock is a transport limitation rather than a missing capability: Anthropic
exposes server-side tool search there only through the InvokeModel API, and
the gateway routes Bedrock through the Converse API.
### Gateway Response Cache [#gateway-response-cache]
If [gateway caching](https://docs.vichar.io/features/caching/gateway-caching) is enabled on the project, a repeated request with the same [cache-key fields](https://docs.vichar.io/features/caching/gateway-caching#cache-key-generation) — resolved provider and model, messages, and the other keyed parameters; fields outside the cache key and insignificant JSON whitespace don't affect matching, but the key order inside message and tool objects does, and a request routed to a different provider is a cache miss — is replayed from cache instead of being sent upstream. Because this endpoint exposes no metadata envelope or cost fields, the replayed body is indistinguishable from the original (same `id`, content, and token counts), so the `x-llmgateway-cache: HIT` response header is the marker to check. Send `x-no-cache: true` to bypass the cache for a single request.
# Vichar Authentication, API Keys & IAM Rules
URL: https://docs.vichar.io/features/api-keys
API keys are the primary method for authenticating with the Vichar. This guide covers creating API keys, managing them, and configuring IAM rules for fine-grained access control.
## Overview [#overview]
Vichar provides comprehensive API key management with the following features:
* **Basic API Key Management**: Create, list, rename, update, and delete API keys
* **Usage Limits**: Set lifetime and recurring spending limits on individual API keys
* **Expiration (TTL)**: Give a key a time-to-live so it disables itself automatically
* **Rotation (Rolling)**: Replace a key's secret in place without losing its settings or history
* **IAM Rules**: Fine-grained access control for models, providers, and pricing
* **Usage Tracking**: Monitor API key usage and costs
* **Status Management**: Enable/disable keys without deletion
This page covers gateway API keys (`vichar_…`), the keys you send to the
gateway as a bearer token.
## Creating API Keys [#creating-api-keys]
### Via Dashboard [#via-dashboard]
1. Navigate to your project in the Vichar dashboard
2. Go to the **API Keys** section
3. Click **Create API Key**
4. Provide a description for your key
5. Optionally set an all-time usage limit
6. Optionally set a recurring usage limit such as `$10 / day` or `$500 / month`
7. Optionally set an expiration (TTL) such as `30 minutes`, `12 hours`, or `7 days`
8. Click **Create**
API keys are shown in full only once during creation. Make sure to copy and
store them securely. New and rolled secrets are stored only as keyed
HMAC-SHA-256 fingerprints, and authentication compares the fingerprint of the
presented secret.
## Using API Keys [#using-api-keys]
Once you have an API key, use it in the `Authorization` header of your requests:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer vichar_your_api_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
## Renaming API Keys [#renaming-api-keys]
A key's name is only a label, so you can change it at any time from the
dashboard without affecting the secret, its usage history, limits, or IAM rules.
## Budget alerts [#budget-alerts]
Enable **API key budgets** in the dashboard's notification settings to receive in-app or email warnings as a key approaches its lifetime or recurring limit. Choose a threshold from 50% to 100%; the default is 80%. See [Notifications](https://docs.vichar.io/features/notifications) for delivery and access rules.
## Disabling/Enabling API Keys [#disablingenabling-api-keys]
You can disable an API key to stop it from being used, but the key is not deleted and can be re-enabled later.
## Rotating (Rolling) API Keys [#rotating-rolling-api-keys]
Rolling a key generates a **new secret for the same key** and invalidates the old
one immediately. Everything else about the key is preserved: its name, usage
history and statistics, all-time and recurring limits (including the active
period window), IAM rules, and expiration.
Use this when a secret may have been exposed — in a commit, a log, a CI
artifact, or a shared environment — and you want to cut off the leaked value
without losing the key's spend tracking or access rules.
1. Open the **API Keys** page and pick the key's actions menu
2. Choose **Roll Key** and confirm
3. Copy the new secret and update every client that used the old one
The old secret stops working the moment the key is rolled, and the new secret
is shown only once. Roll during a window where you can update your clients
promptly.
Requests made with the old secret are rejected with a `401 Unauthorized`.
## Expiration (TTL) [#expiration-ttl]
You can give an API key a **time-to-live (TTL)** when you create it. Set how long
the key should live — in **minutes**, **hours**, or **days** — and it will be
disabled automatically once that time passes. This is ideal for short-lived
integrations, demos, CI jobs, and temporary access.
* A key works normally until its expiration time
* Once expired, the gateway rejects requests with that key with a `401 Unauthorized`
* A background job marks expired keys as **inactive**, so the dashboard reflects
the disabled state
* Keys created without a TTL never expire (the default)
### Reactivating an Expired Key [#reactivating-an-expired-key]
An expired key is paused, not deleted. To bring it back online you must reactivate
it **with a new future expiration** — an expired key cannot be re-enabled while its
TTL is still in the past. Keys that have no TTL, or whose TTL is still in the
future, can be enabled and disabled freely without setting a new expiration.
Expiration is independent of usage limits. A key can hit its TTL before, or
instead of, reaching a spend cap.
## Usage Limits [#usage-limits]
Usage is tracked per API key on the API Keys page, giving you complete
visibility into spend per key.
You can set two independent limits for each key:
* **All-time usage limit**: A lifetime spend cap
* **Recurring usage limit**: A spend cap that resets every configured hour, day, week, or month
When a key reaches either limit, requests using that key return `401
Unauthorized` until the key is updated or, for recurring limits, the next
usage window starts. This is separate from IAM rule violations, which return
`403 Forbidden`.
Recurring windows support:
* Minimum duration: **1 hour**
* Maximum duration: **12 months**
* Units: **hour**, **day**, **week**, **month**
For the dashboard walkthrough and field-by-field details, see [API Keys in
Learn](https://docs.vichar.io/learn/api-keys).
## IAM Rules [#iam-rules]
IAM (Identity Access Management) rules provide fine-grained access control over what models, providers, and pricing tiers an API key can access.
### Rule Types [#rule-types]
#### Model Access Rules [#model-access-rules]
Control access to specific models:
* **Allow Models**: Only allow access to specific models
* **Deny Models**: Block access to specific models
#### Provider Access Rules [#provider-access-rules]
Control access to specific providers:
* **Allow Providers**: Only allow access to specific providers
* **Deny Providers**: Block access to specific providers
#### Pricing Rules [#pricing-rules]
Control access based on model pricing:
* **Allow Pricing**: Set constraints on what pricing tiers are allowed
* **Deny Pricing**: Block specific pricing tiers
* **Free vs Paid**: Allow or deny access to free vs paid models
#### IP Address Rules [#ip-address-rules]
IP address rules are available on the **Enterprise** plan only. Contact us at
[contact@vichar.io](mailto:contact@vichar.io) to enable them for your organization.
Restrict where the API key can be used from by source IP, using CIDR ranges:
* **Allow IP Ranges (CIDR)**: Only permit requests from the listed IPv4/IPv6 CIDRs
* **Deny IP Ranges (CIDR)**: Block requests from the listed IPv4/IPv6 CIDRs
Both IPv4 (e.g. `192.0.2.0/24`) and IPv6 (e.g. `2001:db8::/32`) ranges are supported, and you can mix both in a single rule. To restrict to a single address, use a `/32` (IPv4) or `/128` (IPv6) prefix.
The gateway reads the client IP from the first entry in the `X-Forwarded-For` header (set by the GCP load balancer). When an `allow_ip_cidrs` rule is configured and the gateway cannot determine the client IP, the request is denied. Invalid CIDR syntax is rejected at rule-creation time with a `400` error.
### Combining Multiple Rules [#combining-multiple-rules]
* **Allow rules of the same type are unioned**: a request passes if it matches *any* of them. For example, one `allow_models` rule with `["claude-opus-4-6"]` and another with `["claude-fable-5"]` allow both models — exactly as if you had a single rule listing both.
* **Allow rules of different types are combined with AND**: the request must satisfy every configured allow rule type (e.g. the model must be in the allowed models *and* served by an allowed provider).
* **Deny rules always apply**: a request matching any deny rule is rejected, regardless of allow rules.
* **Moderation is special-cased**: `/v1/moderations` runs a fixed moderation model that cannot appear in model allowlists, so only provider and IP rules apply to it — model and pricing rules are skipped.
## Member-Level IAM Rules [#member-level-iam-rules]
The same rule types can also be configured **per organization member** by owners and admins on the [Team page](https://docs.vichar.io/learn/team) (admins cannot modify an owner's rules). Member-level rules act as an organization-wide ceiling for that member:
* A request must pass **both** the member's rules and the API key's rules. Within each level, rules combine exactly as described above.
* Key rules can only **narrow** access further — they can never grant anything the member's rules deny. For example, if an admin restricts a member to a single approved [provider](https://app.vichar.io/dashboard) with an `allow_providers` rule, the member can create a key rule allowing only a specific model from that provider, but a key rule allowing any other provider has no effect.
* A key with **no rules of its own** is still fully constrained by its owner's member-level rules.
* Member-level rules apply to all regular API keys created by that member, across every project in the organization.
When a request is denied by a member-level rule, the `403` error message states that the restriction is an organization member IAM rule set by the org admin (rather than the key's own IAM configuration), so key holders know who to contact.
## Team-Level IAM Rules [#team-level-iam-rules]
The same rule types can additionally be configured on an **organization team**. Team rules apply to members with the `developer` role who are assigned to that team, and they are evaluated **before** member and key rules: a request must pass the team's rules, then the member's rules, then the key's rules, and each layer can only narrow what the previous one allowed. When a request is denied by a team-level rule, the `403` error message states that the restriction is inherited from an organization team IAM rule set by the org admin.
## Error Handling [#error-handling]
When API keys encounter IAM rule violations, the API returns a `403` with the standard OpenAI error envelope:
```json
{
"error": {
"message": "Access denied: Model gpt-4 is not in the allowed models list",
"type": "invalid_request_error",
"param": null,
"code": "permission_denied"
}
}
```
Common error scenarios:
* Model not allowed by IAM rules
* Provider blocked by IAM rules
* Pricing limits exceeded
* API key disabled or deleted
* API key expired (TTL passed)
* API key rolled, so the old secret is no longer valid
* Usage limit reached
## Migration from Legacy Keys [#migration-from-legacy-keys]
If you have existing API keys without IAM rules:
1. **Backward Compatibility**: Existing keys continue to work without restrictions
2. **Gradual Migration**: Add IAM rules incrementally
3. **Testing**: Test IAM rules in development before applying to production
4. **Monitoring**: Monitor for access denied errors after implementing rules
API keys without IAM rules have unrestricted access to all models and
providers.
# Cost Breakdown
URL: https://docs.vichar.io/features/cost-breakdown
Vichar provides real-time cost information for each API request directly in the response's `usage` object. This allows you to track costs programmatically without needing to query the dashboard.
Cost breakdown is available for all users on both hosted and self-hosted
deployments.
## Response Format [#response-format]
API responses include cost fields in the `usage` object:
```json
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1234567890,
"model": "openai/gpt-4o",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 15,
"total_tokens": 25,
"cost": 0.000125,
"cost_details": {
"upstream_inference_cost": 0.000125,
"upstream_inference_prompt_cost": 0.000025,
"upstream_inference_completions_cost": 0.0001,
"total_cost": 0.000125,
"input_cost": 0.000025,
"output_cost": 0.0001,
"cached_input_cost": 0,
"request_cost": 0,
"web_search_cost": 0,
"image_input_cost": null,
"image_output_cost": null,
"data_storage_cost": 0.00000025
},
"prompt_tokens_details": {
"cached_tokens": 0,
"cache_write_tokens": 0,
"audio_tokens": 0,
"video_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 0,
"image_tokens": 0,
"audio_tokens": 0
}
}
}
```
## Cost Fields [#cost-fields]
| Field | Description |
| -------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `cost` | Total inference cost for the request in USD |
| `cost_details.upstream_inference_cost` | Combined upstream inference cost in USD (prompt + completions) |
| `cost_details.upstream_inference_prompt_cost` | Upstream cost for prompt tokens in USD (includes cached prompt discount) |
| `cost_details.upstream_inference_completions_cost` | Upstream cost for completion tokens in USD |
| `cost_details.total_cost` | Total request cost in USD (Vichar extended field) |
| `cost_details.input_cost` | Cost for non-cached prompt tokens in USD |
| `cost_details.output_cost` | Cost for completion tokens in USD |
| `cost_details.cached_input_cost` | Cost for cached prompt tokens in USD |
| `cost_details.cache_write_input_cost` | Cost for prompt tokens written to the provider cache in USD, billed at the provider's cache-write premium (e.g. 1.25x for 5m / 2x for 1h on Anthropic) |
| `cost_details.request_cost` | Per-request flat fee in USD (when the model applies one) |
| `cost_details.web_search_cost` | Cost for web search tool calls in USD |
| `cost_details.image_input_cost` | Cost for image inputs in USD |
| `cost_details.image_output_cost` | Cost for image outputs in USD |
| `cost_details.data_storage_cost` | Storage cost for retained request/response payloads in USD |
## Token Detail Fields [#token-detail-fields]
The `usage` object also includes detailed token counters that mirror OpenAI's extended format:
| Field | Description |
| --------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| `prompt_tokens_details.cached_tokens` | Number of prompt tokens served from the provider's prompt cache |
| `prompt_tokens_details.cache_write_tokens` | Number of prompt tokens written into the provider's prompt cache |
| `prompt_tokens_details.cache_creation_tokens` | Alias of `cache_write_tokens`, matching Anthropic's naming |
| `prompt_tokens_details.cache_creation` | Per-TTL breakdown of cache writes (`ephemeral_5m_input_tokens` / `ephemeral_1h_input_tokens`), present when a cache write occurred |
| `prompt_tokens_details.audio_tokens` | Number of audio prompt tokens |
| `prompt_tokens_details.video_tokens` | Number of video prompt tokens |
| `completion_tokens_details.reasoning_tokens` | Number of reasoning tokens produced by reasoning models |
| `completion_tokens_details.image_tokens` | Number of image tokens produced |
| `completion_tokens_details.audio_tokens` | Number of audio tokens produced |
## Streaming Responses [#streaming-responses]
Cost information is also available in streaming responses. The cost fields are included in the final usage chunk sent before the `[DONE]` message:
```
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","choices":[...],"usage":{"prompt_tokens":10,"completion_tokens":15,"total_tokens":25,"cost":0.000125,"cost_details":{"upstream_inference_cost":0.000125,"upstream_inference_prompt_cost":0.000025,"upstream_inference_completions_cost":0.0001,"total_cost":0.000125,"input_cost":0.000025,"output_cost":0.0001,"cached_input_cost":0,"request_cost":0,"web_search_cost":0,"image_input_cost":null,"image_output_cost":null,"data_storage_cost":0.00000025}}}
data: [DONE]
```
## Example: Tracking Costs in Code [#example-tracking-costs-in-code]
Here's an example of how to track costs programmatically using the cost breakdown feature:
```typescript
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LLM_GATEWAY_API_KEY,
baseURL: "https://api.vichar.io/v1",
});
async function trackCosts() {
const response = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Hello!" }],
});
const usage = response.usage as any;
if (usage.cost !== undefined) {
console.log(`Request cost: $${usage.cost.toFixed(6)}`);
console.log(
` Prompt: $${usage.cost_details.upstream_inference_prompt_cost.toFixed(6)}`,
);
console.log(
` Completions: $${usage.cost_details.upstream_inference_completions_cost.toFixed(6)}`,
);
const cachedTokens = usage.prompt_tokens_details?.cached_tokens ?? 0;
if (cachedTokens > 0) {
console.log(` Cached prompt tokens: ${cachedTokens}`);
}
}
return response;
}
```
## Use Cases [#use-cases]
### Budget Monitoring [#budget-monitoring]
Track costs in real-time and implement budget limits in your application:
```typescript
let totalSpent = 0;
const BUDGET_LIMIT = 10.0; // $10 budget
async function makeRequest(messages: Message[]) {
const response = await client.chat.completions.create({
model: "gpt-4o",
messages,
});
const cost = (response.usage as any).cost || 0;
totalSpent += cost;
if (totalSpent > BUDGET_LIMIT) {
throw new Error(`Budget exceeded: $${totalSpent.toFixed(2)}`);
}
return response;
}
```
### Per-User Cost Allocation [#per-user-cost-allocation]
Track costs per user for billing or analytics:
```typescript
const userCosts: Map = new Map();
async function makeRequestForUser(userId: string, messages: Message[]) {
const response = await client.chat.completions.create({
model: "gpt-4o",
messages,
});
const cost = (response.usage as any).cost || 0;
const currentCost = userCosts.get(userId) || 0;
userCosts.set(userId, currentCost + cost);
return response;
}
```
### Cost Analytics [#cost-analytics]
Aggregate costs by model, time period, or any other dimension:
```typescript
interface CostEntry {
timestamp: Date;
model: string;
promptCost: number;
completionsCost: number;
totalCost: number;
}
const costLog: CostEntry[] = [];
async function loggedRequest(model: string, messages: Message[]) {
const response = await client.chat.completions.create({
model,
messages,
});
const usage = response.usage as any;
costLog.push({
timestamp: new Date(),
model: response.model,
promptCost: usage.cost_details?.upstream_inference_prompt_cost || 0,
completionsCost:
usage.cost_details?.upstream_inference_completions_cost || 0,
totalCost: usage.cost || 0,
});
return response;
}
```
## Self-Hosted Deployments [#self-hosted-deployments]
If you're running a self-hosted Vichar deployment, cost breakdown is always included in API responses regardless of plan. This allows you to track internal costs and allocate them across teams or projects.
# Data Retention
URL: https://docs.vichar.io/features/data-retention
Vichar offers configurable data retention policies that allow you to store full request and response payloads. This enables powerful debugging capabilities, detailed analytics, and compliance with data governance requirements.
## Retention Levels [#retention-levels]
Vichar supports two retention levels that can be configured per organization:
| Level | Description | Storage Cost |
| ------------------- | ---------------------------------------------------------------------------------------------- | --------------- |
| **Metadata Only** | Stores request metadata (timestamps, model, tokens, costs) without full payloads. Default. | Free |
| **Retain All Data** | Stores complete request and response payloads including messages, tool calls, and attachments. | $0.01/1M tokens |
Metadata-only retention is enabled by default and provides usage analytics
without additional storage costs.
Retention levels are configurable on standard (pay-as-you-go) organizations
only.
## Storage Pricing [#storage-pricing]
When full data retention is enabled, storage is billed at **$0.01 per 1 million tokens**. This rate applies to:
* Input tokens (prompt)
* Cached input tokens
* Output tokens (completion)
* Reasoning tokens
Storage costs are calculated per request and billed separately from inference. When "Retain All Data" is enabled, each response's `usage.cost_details` object includes a `data_storage_cost` field with the per-request storage cost in USD. See [Cost Breakdown](https://docs.vichar.io/features/cost-breakdown) for the full list of cost fields.
### Example Cost Calculation [#example-cost-calculation]
For a request with:
* 1,000 input tokens
* 500 output tokens
* 1,500 total tokens
Storage cost = 1,500 / 1,000,000 × $0.01 = **$0.000015**
## Configuring Retention [#configuring-retention]
Data retention is configured at the organization level in your dashboard settings. The setting is only available on standard pay-as-you-go organizations:
1. Navigate to **Organization Settings** → **Policies**
2. Select your preferred **Data Retention Level**
3. Save changes
Only organization **owners** can change the retention level — admins and project admins receive a `403`.
Changing retention settings applies to new requests only. Existing stored data
follows the retention period active when it was created.
## Retention Periods [#retention-periods]
Data is retained for 30 days for all users, regardless of plan. After the retention period expires, the stored payloads (prompts, completions, tool payloads, raw request and response bodies) are automatically cleared; request metadata used for analytics is kept.
The Responses API does not require data retention. Stored responses (used for
`previous_response_id` chaining and `GET /v1/responses/:id`) are kept in
dedicated storage for 30 days — matching OpenAI's own retention — regardless
of your organization's data retention policy, and are not billed as data
storage. Send `store: false` with the request to opt out. When an Enterprise
organization's zero-data-retention policy is active, the gateway rejects
Responses API requests unless `store: false` is set, does not retain
compaction state, bypasses the gateway response cache, strips provider
prompt-cache markers, and prevents payload retention or project response
caching from being enabled until ZDR is disabled. ZDR cannot be enabled until
both settings are off.
## Accessing Stored Data [#accessing-stored-data]
When data retention is enabled, you can access your stored requests through the dashboard:
* View request history with full payload inspection
* Filter by model and date range
* Inspect complete request and response payloads
## Use Cases [#use-cases]
### Debugging [#debugging]
Full data retention enables you to:
* Inspect exact prompts sent to models
* Review complete responses including tool calls
* Trace conversation histories
* Identify issues in production
### Analytics [#analytics]
With stored payloads, you can:
* Analyze prompt patterns and effectiveness
* Track response quality over time
* Build custom dashboards and reports
* Measure model performance across use cases
### Compliance [#compliance]
Data retention helps meet compliance requirements by:
* Maintaining audit trails of AI interactions
* Enabling data governance policies
* Supporting incident investigation
* Providing records for regulatory requirements
## Billing Considerations [#billing-considerations]
### Credit Usage [#credit-usage]
In **API keys mode** (using your own provider keys):
* Only storage costs are deducted from Vichar credits
* Inference costs are billed directly to your provider
In **credits mode**:
* Both inference and storage costs are deducted from credits
### Monitoring Storage Costs [#monitoring-storage-costs]
Storage costs appear in:
* Usage dashboard under "Storage" category
* Billing invoices as a separate line item
Enable [auto top-up](https://app.vichar.io/dashboard) in billing settings to
ensure uninterrupted service when storage costs accumulate.
## Self-Hosted Deployments [#self-hosted-deployments]
Self-hosted deployments have full control over data retention:
* Enable or disable the 30-day cleanup job with `ENABLE_DATA_RETENTION_CLEANUP=true` (disabled by default; the period itself is not configurable)
* Data is stored in your own PostgreSQL database
* No additional storage costs (you manage your own infrastructure)
## Privacy and Security [#privacy-and-security]
* All stored data is encrypted at rest
* Access is restricted to organization members with appropriate permissions
* Stored payloads are automatically cleared after the retention period (on self-hosted deployments, only when the cleanup job is enabled via `ENABLE_DATA_RETENTION_CLEANUP=true`)
* You can request immediate deletion of specific records through support
# Document Reading
URL: https://docs.vichar.io/features/documents
Vichar supports sending documents (PDFs and other file types) to document-capable models using OpenAI's `file` content block format. The gateway forwards the document to the underlying provider so the model can read and reason over its contents.
## Document-Capable Models [#document-capable-models]
Document input is currently supported on Google Gemini models via Google AI Studio. You can find document-capable models on the [models page with the document filter](https://app.vichar.io/dashboard).
## Sending a Document [#sending-a-document]
Add a `file` content block to a user message. The `file_data` field must be a base64-encoded data URL that includes the document's MIME type.
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.6-flash",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Summarize this document."
},
{
"type": "file",
"file": {
"filename": "report.pdf",
"file_data": "data:application/pdf;base64,JVBERi0xLjQKJ..."
}
}
]
}
]
}'
```
### Content Block Fields [#content-block-fields]
* **`type`**: must be `"file"`.
* **`file.filename`** *(optional)*: original filename, forwarded for context.
* **`file.file_data`**: base64-encoded data URL of the form `data:;base64,`.
The `file.file_id` field (for referencing files uploaded via a provider's
Files API) is accepted by the schema but not currently supported by the Google
transform. Use `file_data` with an inline base64 data URL.
## Supported File Types [#supported-file-types]
The accepted MIME types depend on the target model. Gemini models commonly support:
* `application/pdf`
* `text/plain`
* `text/html`
* `text/css`
* `text/javascript`
* `text/csv`
* `text/markdown`
* `text/xml`
If the upstream provider rejects the MIME type, the gateway surfaces a `400` error including the unsupported MIME type and the provider it was sent to. To use a different file type, encode the file with the matching MIME type in the data URL prefix.
## Encoding a File as a Data URL [#encoding-a-file-as-a-data-url]
Any tool that can produce base64 output works. For example, in a shell:
```bash
DATA=$(base64 -i report.pdf | tr -d '\n')
echo "data:application/pdf;base64,$DATA"
```
Or in JavaScript:
```javascript
import { readFileSync } from "node:fs";
const buffer = readFileSync("report.pdf");
const fileData = `data:application/pdf;base64,${buffer.toString("base64")}`;
```
Then pass `fileData` as the `file.file_data` value in your request.
## Multiple Documents [#multiple-documents]
You can include multiple `file` blocks in a single message, optionally mixed with text and image content:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.1-pro-preview",
"messages": [
{
"role": "user",
"content": [
{ "type": "text", "text": "Compare these two reports." },
{
"type": "file",
"file": {
"filename": "q1.pdf",
"file_data": "data:application/pdf;base64,JVBERi0x..."
}
},
{
"type": "file",
"file": {
"filename": "q2.pdf",
"file_data": "data:application/pdf;base64,JVBERi0x..."
}
}
]
}
]
}'
```
## Error Handling [#error-handling]
The gateway returns `400` for the following document-related errors:
* The selected model does not support document input.
* The `file` block is missing both `file_data` and `file_id`.
* `file_data` is not a valid base64 data URL.
* The upstream provider rejects the document's MIME type for the selected model.
# Embeddings
URL: https://docs.vichar.io/features/embeddings
Vichar exposes an OpenAI-compatible `/v1/embeddings` endpoint for generating vector representations of text — useful for semantic search, clustering, recommendations, and RAG.
Browse available embedding models on the [models page](https://app.vichar.io/dashboard).
## Supported providers [#supported-providers]
* **OpenAI** — `text-embedding-3-small`, `text-embedding-3-large`, `text-embedding-ada-002`
* **Google AI Studio** — `gemini-embedding-2` (recommended), `gemini-embedding-001` (legacy)
* **Google Vertex AI** — `gemini-embedding-001`, `text-embedding-005`
The gateway translates between provider-native request/response shapes (e.g. Google's `:embedContent` / `:batchEmbedContents`) and the OpenAI-compatible payload, so you can swap models without changing your client code.
## cURL [#curl]
```bash
curl -X POST "https://api.vichar.io/v1/embeddings" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-small",
"input": "The quick brown fox jumps over the lazy dog."
}'
```
## OpenAI JS SDK [#openai-js-sdk]
```ts
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LLM_GATEWAY_API_KEY,
baseURL: "https://api.vichar.io/v1",
});
const response = await client.embeddings.create({
model: "text-embedding-3-small",
input: "The quick brown fox jumps over the lazy dog.",
});
console.log(response.data[0].embedding);
```
Embedding models are billed only for input tokens. There are no output tokens
since embeddings are fixed-size vectors.
## Improving retrieval quality [#improving-retrieval-quality]
Vector similarity is fast but approximate. For RAG, pair embeddings with
[rerank](https://docs.vichar.io/features/rerank): use embeddings to pull a broad candidate set, then
rerank those candidates to pick the few documents that actually go in the
prompt.
# Image Generation
URL: https://docs.vichar.io/features/image-generation
Vichar supports image generation through two APIs:
1. **`/v1/images/generations`** — OpenAI-compatible images endpoint (recommended for simple image generation)
2. **`/v1/images/edits`** — OpenAI-compatible image editing endpoint
3. **`/v1/chat/completions`** — Chat completions with image generation models (for conversational image generation and editing)
For asynchronous video generation, see [Video Generation](https://docs.vichar.io/features/video-generation).
## Available Models [#available-models]
You can find all available image generation models on our [models page](https://app.vichar.io/dashboard).
## OpenAI Images API [#openai-images-api]
The `/v1/images/generations` endpoint provides a drop-in replacement for OpenAI's image generation API. It works with any OpenAI-compatible client library.
### Parameters [#parameters]
| Parameter | Type | Default | Description |
| ----------------- | ------- | ------------ | ---------------------------------------------------------------------------------------------------------------------------- |
| `prompt` | string | required | A text description of the desired image(s) |
| `model` | string | `"auto"` | The model to use. `auto` resolves to `gemini-3-pro-image` |
| `n` | integer | `1` | Number of images to generate (1-10) |
| `size` | string | — | Image dimensions. Supported sizes depend on the model/provider — see [Image Configuration](#image-configuration) |
| `quality` | string | — | Image quality. Supported values depend on the model/provider — see [Image Configuration](#image-configuration) |
| `moderation` | string | `"auto"` | Content filtering strictness for models that support it: `auto` or `low` — see [Moderation](#moderation) |
| `response_format` | string | `"b64_json"` | Only `b64_json` is supported |
| `style` | string | — | Image style: `vivid` or `natural` |
| `service_tier` | string | — | Processing tier for mappings that offer one: `flex`, `priority`, or `default` — see [Service Tiers](https://docs.vichar.io/features/service-tiers) |
### curl [#curl]
```bash
curl -X POST "https://api.vichar.io/v1/images/generations" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-pro-image",
"prompt": "A cute cat wearing a tiny top hat",
"n": 1,
"size": "1024x1024"
}'
```
### Usage and cost [#usage-and-cost]
Both `/v1/images/generations` and `/v1/images/edits` return a `usage` object with the token counts and the billed cost of the request, in USD:
```json
"usage": {
"input_tokens": 10,
"input_tokens_details": { "image_tokens": 0, "text_tokens": 10 },
"output_tokens": 229,
"output_tokens_details": { "image_tokens": 229, "text_tokens": 0 },
"total_tokens": 239,
"cost": 0.00692,
"cost_details": { "input_cost": 0.00005, "image_output_cost": 0.00687 }
}
```
`cost_details` has the same fields as `usage.cost_details` on [chat completions](#response-format).
### OpenAI SDK [#openai-sdk]
Works with the standard OpenAI client library — just point the base URL to Vichar.
```ts
import OpenAI from "openai";
import { writeFileSync } from "fs";
const client = new OpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const response = await client.images.generate({
model: "gemini-3-pro-image",
prompt: "A futuristic city skyline at sunset with flying cars",
n: 1,
size: "1024x1024",
});
response.data.forEach((image, i) => {
if (image.b64_json) {
const buf = Buffer.from(image.b64_json, "base64");
writeFileSync(`image-${i}.png`, buf);
}
});
```
### Vercel AI SDK [#vercel-ai-sdk]
Use the `@llmgateway/ai-sdk-provider` with `generateImage`.
```ts
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { generateImage } from "ai";
import { writeFileSync } from "fs";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const result = await generateImage({
model: llmgateway.image("gemini-3-pro-image"),
prompt:
"A cozy cabin in a snowy mountain landscape at night with aurora borealis",
size: "1024x1024",
n: 1,
// aspectRatio and quality are model-specific — only some providers honor them.
// aspectRatio works on Gemini image models; OpenAI gpt-image-2 ignores it
// (use a literal WxH `size` instead).
aspectRatio: "16:9",
// quality works on OpenAI gpt-image-2 ("low" | "medium" | "high" | "auto").
// The AI SDK only forwards it through providerOptions.
providerOptions: {
llmgateway: { quality: "high" },
},
});
result.images.forEach((image, i) => {
const buf = Buffer.from(image.base64, "base64");
writeFileSync(`image-${i}.png`, buf);
});
```
## OpenAI Images Edit API [#openai-images-edit-api]
The `/v1/images/edits` endpoint is OpenAI-compatible and supports a focused subset of `images.edit` parameters.
### Parameters [#parameters-1]
| Parameter | Type | Required | Description |
| -------------------- | ------------------------ | -------- | --------------------------------------------------------------------------------- |
| `images` | array of `{ image_url }` | yes | Input images. `image_url` supports HTTPS URLs and base64 data URLs |
| `prompt` | string | yes | A text description of the desired image edit |
| `model` | string | no | Image editing model |
| `background` | enum | no | `transparent`, `opaque`, or `auto` |
| `input_fidelity` | enum | no | `high` or `low` |
| `n` | integer | no | Number of edited images to generate |
| `output_format` | enum | no | `png`, `jpeg`, or `webp` |
| `output_compression` | integer | no | Compression level for `jpeg`/`webp` |
| `quality` | enum | no | `low`, `medium`, `high`, or `auto`; GPT Image 2.5 also supports `xhigh` and `max` |
| `moderation` | enum | no | `auto` or `low` — see [Moderation](#moderation) |
| `size` | string | no | Output size. Examples: `1024x1024`, `1536x1024`, `1K`, `2K`, `4K` |
| `aspect_ratio` | string | no | Aspect ratio override. Examples: `1:1`, `16:9`, `4:3`, `5:4` |
| `service_tier` | enum | no | `flex`, `priority`, or `default` — see [Service Tiers](https://docs.vichar.io/features/service-tiers) |
`mask` is not supported yet on `/v1/images/edits`.
### curl (HTTPS image URL) [#curl-https-image-url]
```bash
curl -X POST "https://api.vichar.io/v1/images/edits" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"images": [
{
"image_url": "https://example.com/source-image.png"
}
],
"prompt": "Add a watercolor effect to this image",
"model": "gemini-3-pro-image",
"aspect_ratio": "16:9",
"quality": "high",
"size": "4K"
}'
```
### curl (base64 data URL) [#curl-base64-data-url]
```bash
curl -X POST "https://api.vichar.io/v1/images/edits" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"images": [
{
"image_url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAA..."
}
],
"prompt": "Turn this into a pixel-art style image"
}'
```
## Chat Completions API [#chat-completions-api]
Image generation also works through the `/v1/chat/completions` endpoint, which is useful for conversational image generation, image editing with vision, and multi-turn interactions.
### Making Requests [#making-requests]
Simply use an image generation model and provide a text prompt describing the image you want to create.
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-pro-image",
"messages": [
{
"role": "user",
"content": "Generate an image of a cute golden retriever puppy playing in a sunny meadow"
}
]
}'
```
### Response Format [#response-format]
Image generation models return responses in the standard chat completions format, with generated images included in the `images` array within the assistant message:
```json
{
"id": "chatcmpl-1756234109285",
"object": "chat.completion",
"created": 1756234109,
"model": "gemini-3-pro-image",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Here's an image of a cute dog for you: ",
"images": [
{
"type": "image_url",
"image_url": {
"url": "data:image/png;base64,"
}
}
]
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 8,
"completion_tokens": 1303,
"total_tokens": 1311
}
}
```
A request stopped by the [gateway content filter](https://docs.vichar.io/resources/error-handling#gateway-content-filter) returns `finish_reason: "content_filter"` on the Chat Completions API and an empty `data` array on the Images API.
A provider's own safety rejection is returned the same way, and it is billed the way the provider bills us: a rejection the provider returns as an error carries no usage and costs nothing, apart from a rejection fee the provider publishes (currently only xAI's chat models). A block the provider serves as a normal response with usage, as Gemini image models do, is charged for the input tokens it reports, never for the blocked image.
### Vision support [#vision-support]
You can edit or modify images by combining image generation with [vision models](https://docs.vichar.io/features/vision) by including the image in the `messages` array.
### Response Structure [#response-structure]
#### Images Array [#images-array]
The `images` array contains one or more generated images with the following structure:
* `type`: Always `"image_url"` for generated images
* `image_url.url`: A data URL containing the base64-encoded image data (format: `data:image/png;base64,`)
#### Content Field [#content-field]
The `content` field may contain descriptive text about the generated image, depending on the model's behavior.
### AI SDK (Chat Completions) [#ai-sdk-chat-completions]
You can use the AI SDK to generate images with your existing generateText or streamText calls using the Vichar provider.
#### Example [#example]
```ts title="/api/chat/route.ts"
import { streamText, type UIMessage, convertToModelMessages } from "ai";
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
interface ChatRequestBody {
messages: UIMessage[];
}
export async function POST(req: Request) {
const body = await req.json();
const { messages }: ChatRequestBody = body;
const llmgateway = createLLMGateway({
apiKey: "llmgateway_api_key",
baseUrl: "https://api.vichar.io/v1",
});
try {
const result = streamText({
model: llmgateway.chat("gemini-3-pro-image"),
messages: convertToModelMessages(messages),
});
return result.toUIMessageStreamResponse();
} catch {
return new Response(JSON.stringify({ error: "Vichar request failed" }), {
status: 500,
});
}
}
```
Then you can render the image in your frontend using the `Image` component from the [ai-elements](https://ai-sdk.dev/elements/components/image).
Here is a full example of how to use the AI SDK to generate images in your frontend:
```tsx title="/app/page.tsx"
"use client";
import { useState, useRef } from "react";
import { useChat } from "@ai-sdk/react";
import { parseImagePartToDataUrl } from "@/lib/image-utils";
import {
PromptInput,
PromptInputBody,
PromptInputButton,
PromptInputSubmit,
PromptInputTextarea,
PromptInputToolbar,
} from "@/components/ai-elements/prompt-input";
import {
Conversation,
ConversationContent,
} from "@/components/ai-elements/conversation";
import { Image } from "@/components/ai-elements/image";
import { Loader } from "@/components/ai-elements/loader";
import { Message, MessageContent } from "@/components/ai-elements/message";
import { Response } from "@/components/ai-elements/response";
export const ChatUI = () => {
const textareaRef = useRef(null);
const [text, setText] = useState("");
const { messages, status, stop, regenerate, sendMessage } = useChat();
return (
<>
>
);
};
```
```ts title="/lib/image-utils.ts"
/**
* Parses a file object containing image data and returns a properly formatted data URL
* and normalized media type.
*
* Handles:
* - Normalizing mediaType from various property names (mediaType, mime_type)
* - Detecting existing data: URLs
* - Detecting base64-looking content
* - Stripping whitespace from base64 content
* - Building proper data:...;base64,... URLs
*/
export function parseImageFile(file: {
url?: string;
mediaType?: string;
mime_type?: string;
}): { dataUrl: string; mediaType: string } {
const mediaType = file.mediaType || file.mime_type || "image/png";
let url = String(file.url || "");
const isDataUrl = url.startsWith("data:");
const looksLikeBase64 =
!isDataUrl && /^[A-Za-z0-9+/=\s]+$/.test(url.slice(0, 200));
if (looksLikeBase64) {
url = url.replace(/\s+/g, "");
}
const dataUrl = isDataUrl
? url
: looksLikeBase64
? `data:${mediaType};base64,${url}`
: url;
return { dataUrl, mediaType };
}
/**
* Extracts base64-only content from a data URL.
* Returns empty string if the input is not a valid data URL.
*/
export function extractBase64FromDataUrl(dataUrl: string): string {
if (!dataUrl.startsWith("data:")) {
return "";
}
const comma = dataUrl.indexOf(",");
return comma >= 0 ? dataUrl.slice(comma + 1) : "";
}
/**
* Parses an image part (either image_url or file type) and returns
* dataUrl, base64Only, and mediaType ready for rendering.
*
* Handles error cases gracefully by returning empty base64Only string
* when parsing fails, allowing the renderer to skip invalid images.
*/
export function parseImagePartToDataUrl(part: any): {
dataUrl: string;
base64Only: string;
mediaType: string;
} {
try {
// Handle image_url parts
if (part.type === "image_url" && part.image_url?.url) {
const url = part.image_url.url;
const mediaType = "image/png"; // Default for image_url parts
if (url.startsWith("data:")) {
// Extract media type from data URL if present
const match = url.match(/data:([^;]+)/);
const extractedMediaType = match?.[1] || mediaType;
return {
dataUrl: url,
base64Only: extractBase64FromDataUrl(url),
mediaType: extractedMediaType,
};
}
return {
dataUrl: url,
base64Only: "",
mediaType,
};
}
// Handle file parts (AI SDK format)
if (part.type === "file") {
const { dataUrl, mediaType } = parseImageFile(part);
return {
dataUrl,
base64Only: extractBase64FromDataUrl(dataUrl),
mediaType,
};
}
return {
dataUrl: "",
base64Only: "",
mediaType: "image/png",
};
} catch {
return {
dataUrl: "",
base64Only: "",
mediaType: "image/png",
};
}
}
```
## Image Configuration [#image-configuration]
You can customize the generated image using the optional `image_config` parameter (for chat completions) or `size`/`quality`/`style` parameters (for the images API). The supported parameters vary by provider.
### Google Models [#google-models]
Available Google models:
| Model | Description |
| ------------------------ | ----------------------------------------------------------------------------------- |
| `gemini-3-pro-image` | Gemini 3 Pro with native image generation. Supports aspect ratios and 1K–4K sizes. |
| `gemini-3.1-flash-image` | Gemini 3.1 Flash with native image generation. Supports 0.5K–4K sizes (default 1K). |
#### gemini-3-pro-image [#gemini-3-pro-image]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-pro-image",
"messages": [
{
"role": "user",
"content": "Generate an image of a mountain landscape at sunset"
}
],
"image_config": {
"aspect_ratio": "16:9",
"image_size": "4K"
}
}'
```
| Parameter | Type | Description |
| -------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `aspect_ratio` | string | The aspect ratio of the generated image. Options: `"1:1"`, `"2:3"`, `"3:2"`, `"3:4"`, `"4:3"`, `"4:5"`, `"5:4"`, `"9:16"`, `"16:9"`, `"21:9"` |
| `image_size` | string | The resolution of the generated image. Options: `"1K"` (1024x1024), `"2K"` (2048x2048), `"4K"` (4096x4096) |
#### gemini-3.1-flash-image [#gemini-31-flash-image]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.1-flash-image",
"messages": [
{
"role": "user",
"content": "Generate an image of a mountain landscape at sunset"
}
],
"image_config": {
"image_size": "1K"
}
}'
```
| Parameter | Type | Description |
| -------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `aspect_ratio` | string | The aspect ratio of the generated image. Options: `"1:1"`, `"1:4"`, `"1:8"`, `"2:3"`, `"3:2"`, `"3:4"`, `"4:1"`, `"4:3"`, `"4:5"`, `"5:4"`, `"8:1"`, `"9:16"`, `"16:9"`, `"21:9"` |
| `image_size` | string | The resolution of the generated image. Options: `"0.5K"` (512x512), `"1K"` (1024x1024, default), `"2K"` (2048x2048), `"4K"` (4096x4096) |
`gemini-3.1-flash-image` uniquely supports `"0.5K"` resolution, which is not
available on other Google image models.
### Meta Models [#meta-models]
Muse Image uses Meta's Responses API for generation and editing. It reasons
before rendering and can use reference images across refinement turns.
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "meta/muse-image-1.0",
"messages": [
{
"role": "user",
"content": "Create a product photo on a warm studio background"
}
],
"image_config": {
"image_size": "1024x1536"
}
}'
```
| Parameter | Type | Description |
| ------------ | ------ | -------------------------------------------------------------------------- |
| `image_size` | string | One of `"1024x1024"`, `"1024x1536"`, or `"1536x1024"`. Defaults to square. |
Muse Image does not expose a quality setting. Use `image_size` to choose square,
portrait, or landscape output.
### Alibaba Models [#alibaba-models]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/qwen-image-3.0",
"messages": [
{
"role": "user",
"content": "Generate an image of a mountain landscape at sunset"
}
],
"image_config": {
"image_size": "1024x1536",
"n": 1,
"seed": 42
}
}'
```
| Parameter | Type | Description |
| ------------ | ------- | ------------------------------------------------------------------------------------------------ |
| `image_size` | string | Image dimensions in `WIDTHxHEIGHT` format. Examples: `"1024x1024"`, `"1024x1536"`, `"1536x1024"` |
| `n` | integer | Number of images to generate (1-4) |
| `seed` | integer | Random seed for reproducible generation |
Available Alibaba models (see the [models page](https://app.vichar.io/dashboard) for current pricing):
| Model | Description |
| ---------------------------- | ------------------------------------------------------------------------------------ |
| `alibaba/qwen-image-3.0` | Third-generation image generation and editing |
| `alibaba/qwen-image-3.0-pro` | Highest quality third-generation generation and editing. Priced per output size tier |
Alibaba models use explicit pixel dimensions (e.g., `"1024x1536"`) instead of
aspect ratios. For portrait orientation use `"1024x1536"`, for landscape use
`"1536x1024"`.
### Z.AI Models [#zai-models]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "zai/cogview-4",
"messages": [
{
"role": "user",
"content": "Generate an image of a futuristic city skyline"
}
],
"image_config": {
"image_size": "1024x1024"
}
}'
```
| Parameter | Type | Description |
| ------------ | ------- | ------------------------------------------------------------------------------------------------ |
| `image_size` | string | Image dimensions in `WIDTHxHEIGHT` format. Examples: `"1024x1024"`, `"2048x1024"`, `"1024x2048"` |
| `n` | integer | Number of images to generate |
Available Z.AI models (see the [models page](https://app.vichar.io/dashboard) for current pricing):
| Model | Description |
| --------------- | ------------------------------------------------------------------------------------------------------------------- |
| `zai/cogview-4` | CogView-4 with bilingual support and excellent text rendering |
| `zai/glm-image` | GLM-Image with hybrid auto-regressive architecture, excellent for text-rendering and knowledge-intensive generation |
CogView-4 supports both Chinese and English prompts and excels at generating
images with embedded text.
### OpenAI Models [#openai-models]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2",
"messages": [
{
"role": "user",
"content": "Generate a photo-real cinematic landscape at golden hour"
}
],
"image_config": {
"image_size": "3072x2160",
"image_quality": "low"
}
}'
```
| Parameter | Type | Description |
| --------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------- |
| `image_size` | string | Image dimensions in `WIDTHxHEIGHT` format, or `"auto"` to let the model choose. |
| `image_quality` | string | `"low"`, `"medium"`, `"high"`, or `"auto"`; GPT Image 2.5 also supports `"xhigh"` and `"max"`. Defaults to `"auto"` when omitted. |
| `moderation` | string | `"auto"` or `"low"` — see [Moderation](#moderation). Defaults to `"auto"` when omitted. |
OpenAI image models do **not** accept `aspect_ratio`. Always specify
`image_size` as `WIDTHxHEIGHT` (e.g. `"1024x1024"`, `"3072x2160"`). OpenAI
requires both width and height to be divisible by 16, the longest edge to be ≤
3840, and the total pixel count to fit within the model's pixel budget;
requests outside these bounds are rejected with HTTP 400.
Available OpenAI image models:
| Model | Description |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| `openai/gpt-image-2` | OpenAI's next-generation image model with improved quality and prompt adherence, supporting text and vision. |
| `openai/gpt-image-2.5-sunburst` | Image generation and precise editing with text and image inputs; adds `xhigh` and `max` quality. |
| `openai/gpt-image-2.5-flare` | Fast everyday image generation and editing with text and image inputs; adds `xhigh` and `max` quality. |
GPT Image 2.5 supports `1024x1024`, `1536x1024`, `1024x1536`, `auto`, and
custom sizes within the limits above. Its aspect ratio must stay between 1:3
and 3:1, with 655,360–8,294,400 total pixels. Resolutions above `2560x1440`
are experimental.
Both variants use the same per-token rates as GPT Image 2. Actual image token
usage varies by model, size, quality, and input; billing uses the provider's
reported usage. See the [models page](https://app.vichar.io/dashboard)
for current pricing.
GPT Image 2 and both GPT Image 2.5 variants are also served through Azure at
the same rates — swap the `openai/` prefix for `azure/` to pin that route, or
send the bare model id and let the gateway pick.
### ByteDance Models [#bytedance-models]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bytedance/seedream-4-5",
"messages": [
{
"role": "user",
"content": "Generate an image of a futuristic cyberpunk city at night"
}
],
"image_config": {
"image_size": "2048x2048"
}
}'
```
| Parameter | Type | Description |
| ------------ | ------ | ------------------------------------------------------------------------------------------------ |
| `image_size` | string | Image dimensions in `WIDTHxHEIGHT` format. Examples: `"1024x1024"`, `"2048x2048"`, `"4096x4096"` |
Available ByteDance models (see the [models page](https://app.vichar.io/dashboard) for current pricing):
| Model | Description |
| ---------------------------- | --------------------------------------------------------------- |
| `bytedance/seedream-4-0` | High-quality text-to-image generation with 2K default output |
| `bytedance/seedream-4-5` | Enhanced quality and consistency with improved prompt adherence |
| `bytedance/seedream-5-0-pro` | Precise generation and reference-image editing at 1K or 2K |
Seedream models support up to 2-10 reference images for multi-image fusion and
generation. The default output resolution is 2048×2048 (2K), with support up
to 4096×4096 (4K).
## Moderation [#moderation]
GPT Image models expose a `moderation` parameter that controls how strict the
provider's content filtering is:
* `auto` (default) — standard filtering, which limits certain categories of
potentially age-inappropriate content.
* `low` — less restrictive filtering. OpenAI's content policy still applies.
Send it as a top-level parameter on `/v1/images/generations` and
`/v1/images/edits`, or inside `image_config` on `/v1/chat/completions`:
```bash
curl -X POST "https://api.vichar.io/v1/images/generations" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5-flare",
"prompt": "A cute cat wearing a tiny top hat",
"moderation": "low"
}'
```
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5-flare",
"messages": [{ "role": "user", "content": "A cute cat wearing a tiny top hat" }],
"image_config": { "moderation": "low" }
}'
```
Models without a moderation control ignore the parameter.
## Usage Notes [#usage-notes]
Image generation models typically have higher token costs compared to
text-only models due to the computational requirements of image synthesis.
Generated images are returned as base64-encoded data URLs, which can be large.
Consider the payload size when integrating image generation into your
applications.
# Metadata
URL: https://docs.vichar.io/features/metadata
Vichar supports sending additional metadata with your requests using custom headers. This allows you to include information like user sessions, application versions, tenant IDs, or other contextual data that can be useful for analytics and monitoring.
Later, you can filter by specific values to return, such as for a specific user or session. Additionally, in the future, you will be able to segment your analytics and monitoring based on this metadata. For example, you could show cost and latency breakdowns per user, application, country, feature, or any other dimension you want to track.
## Custom Headers [#custom-headers]
You can include custom headers with the `X-Vichar-` prefix to send metadata alongside your LLM requests:
```bash
curl -X POST https://api.vichar.io/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "X-Vichar-Country: US" \
-H "X-Vichar-User-ID: 9403f741-a524-4b18-b1b2-dbb71cdff2a4" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "Hello, how are you?"
}
]
}'
```
## Best Practices [#best-practices]
### Header Naming [#header-naming]
* Use the `X-Vichar-` prefix for all custom metadata
* Use descriptive, consistent naming conventions
* Avoid special characters; use hyphens to separate words
### Data Privacy [#data-privacy]
* Be mindful of sensitive data in headers
* Consider hashing or anonymizing user identifiers
* Follow your organization's data privacy policies
### Performance [#performance]
* Keep header values reasonably short
* Avoid sending unnecessary metadata that won't be used for analytics
* Consider the impact on request size, especially for high-volume applications
## Example: Multi-tenant Application [#example-multi-tenant-application]
For a multi-tenant application, you might use metadata headers like this:
```bash
curl -X POST https://api.vichar.io/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "X-Vichar-Tenant-ID: acme-corp" \
-H "X-Vichar-User-ID: user-12345" \
-H "X-Vichar-App-Version: 2.1.4" \
-H "X-Vichar-Feature: chat-assistant" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "Summarize this document..."
}
]
}'
```
This allows you to track usage and costs per tenant, user, application version, and feature, providing detailed insights into how your LLM integration is being used across your platform.
# Org Models Directory
URL: https://docs.vichar.io/features/models-directory
The **Models** page in the dashboard (`Organization → Models`) lists every model your organization can route to in a single directory — the full Vichar catalogue, searchable and filterable by capability, provider, price, and context size, exactly like the public [models page](https://app.vichar.io/dashboard).
## Who can see it [#who-can-see-it]
Every active organization member can browse the directory read-only — including project-scoped **developer** members, who see it as their only organization page. This is the recommended way for developers to discover available providers and models without being granted additional permissions.
## Discover models with an API key [#discover-models-with-an-api-key]
Call `GET /v1/models` with your API key to list models available under your
organization's inherited team and member IAM rules, and the key's own IAM rules.
```bash
curl "https://api.vichar.io/v1/models" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY"
```
The endpoint also accepts the `x-api-key` header. Without either authentication
header, it returns the public catalogue. Invalid, inactive, or expired credentials
return `401`.
Set `?include_restricted=true` to return the full public catalogue even with an
authenticated API key. This skips organization, team, member, key, project, and
plan filtering for discovery. Other query filters still apply, and inference
requests still enforce all access restrictions. The default is `false`, which
keeps authenticated results filtered.
Use `?mapped=true` for provider-prefixed entries. Existing query filters, including
`no_training`, apply alongside access restrictions. Discovery reflects model
permissions and project configuration; it does not check remaining credits,
usage budgets, or current provider health.
## Related [#related]
* [Knowledge base → Models](https://docs.vichar.io/learn/models) — a walkthrough of the page itself.
# Moderations
URL: https://docs.vichar.io/features/moderations
Vichar supports the OpenAI-compatible `/v1/moderations` endpoint for text
and multimodal safety classification.
Use it when you want to:
* Screen user prompts before they reach a model
* Review generated output before displaying it
* Apply the same moderation API shape you already use with OpenAI clients
For the full request and response schema, see the
[API reference](https://docs.vichar.io/v1_moderations).
## Endpoint [#endpoint]
`POST https://api.vichar.io/v1/moderations`
Authenticate with your Vichar API key:
```bash
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY"
```
## Supported Inputs [#supported-inputs]
The `input` field accepts:
* A single string
* An array of strings
* An array of multimodal content items with `text` and `image_url`
The default model is `omni-moderation-latest`.
## Pricing [#pricing]
Starting **Friday, August 7, 2026**, moderation requests are billed at a
flat **$0.00001 per request**, independent of the input size, the number of
inputs, or the moderation model you pick. Only successful requests are billed —
failed and retried attempts cost nothing. Before that date, moderation requests
are free.
Because the endpoint becomes paid, it also starts requiring a credit balance
from that date: moderation requests from an organization with no credits are
rejected with `402`. Top up before August 7 to avoid an interruption.
When your project runs in `api-keys` mode and the request is served with your
own OpenAI key, no credits are deducted and no balance is required.
## curl [#curl]
### Single text input [#single-text-input]
```bash
curl -X POST "https://api.vichar.io/v1/moderations" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": "I want to harm someone."
}'
```
### Multiple text inputs [#multiple-text-inputs]
```bash
curl -X POST "https://api.vichar.io/v1/moderations" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "omni-moderation-latest",
"input": [
"This is a harmless sentence.",
"I want to attack somebody."
]
}'
```
### Multimodal input [#multimodal-input]
```bash
curl -X POST "https://api.vichar.io/v1/moderations" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": [
{
"type": "text",
"text": "Check this image for violent content."
},
{
"type": "image_url",
"image_url": {
"url": "https://example.com/image.png"
}
}
]
}'
```
## OpenAI SDK [#openai-sdk]
```ts
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const response = await client.moderations.create({
model: "omni-moderation-latest",
input: "I want to harm someone.",
});
console.log(response.results[0]?.flagged);
```
## Response Shape [#response-shape]
The response follows the standard OpenAI moderation format:
```json
{
"id": "modr-123",
"model": "omni-moderation-latest",
"results": [
{
"flagged": true,
"categories": {
"violence": true,
"self_harm": false
},
"category_scores": {
"violence": 0.98,
"self_harm": 0.01
}
}
]
}
```
## When To Use This Instead Of Chat Content Filtering [#when-to-use-this-instead-of-chat-content-filtering]
Use `/v1/moderations` when you want an explicit moderation decision in your own
application flow.
If you want moderation to happen automatically as part of model requests, use
Vichar content filtering on `/v1/chat/completions` instead.
# Notifications
URL: https://docs.vichar.io/features/notifications
Open the **Notifications** bell in the dashboard header, then select **Notification settings**. Choose **In-app**, **Email**, both, or neither for each alert type. Most alerts start disabled. Email delivery always requires a verified email address. Your preferences apply across projects you can access.
| Alert | When it appears |
| ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| API key budgets | A regular API key reaches your selected percentage of its lifetime or recurring budget. The default threshold is 80%. |
| Model retirements | A model mapping you used in the last 30 days is scheduled for deprecation or deactivation within 30 days. The alert links to the model catalogue so you can choose a newer model. |
| Provider issues | A provider you used in the last 30 days has elevated upstream errors in its recent traffic. |
Budget alerts repeat for each new recurring period, or when you change the limit or alert threshold. Unlimited keys do not trigger budget warnings. Each scheduled model retirement is announced once; ongoing provider issues produce at most one alert per day per project, or per accessible key for developers.
Provider alerts require at least 20 non-cached, non-client-error requests and an upstream error rate of at least 20% in the latest hourly statistics window. Stale statistics do not trigger alerts. Checks run about once a minute; these warnings are not a guarantee that every outage or budget crossing will be caught before a request fails.
Notifications follow your existing project permissions. Developers receive alerts only for their own keys in assigned projects. Removed project access also removes those alerts from the inbox and prevents pending email delivery.
The inbox shows your 50 most recent in-app alerts. Open an alert to follow its action, or select **Mark all as read** to clear the unread indicator. Turn off a channel in notification settings to stop future delivery through it.
# OCR
URL: https://docs.vichar.io/features/ocr
Vichar exposes a dedicated `/v1/ocr` endpoint for optical character
recognition. It extracts text, tables, and layout from PDFs and images and
returns them as clean markdown, one entry per page.
Use it when you want to:
* Turn scanned PDFs or photos into machine-readable markdown
* Pull structured text out of receipts, invoices, forms, or screenshots
* Feed document contents into a downstream model or RAG pipeline
For the full request and response schema, see the
[API reference](https://docs.vichar.io/v1_ocr).
## Endpoint [#endpoint]
`POST https://api.vichar.io/v1/ocr`
Authenticate with your Vichar API key:
```bash
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY"
```
The current model is `mistral-ocr-latest`, billed at **$4 per 1,000 pages**
processed.
## Document Input [#document-input]
The `document` field accepts either a document URL (PDF) or an image:
* `{ "type": "document_url", "document_url": "https://…/file.pdf" }`
* `{ "type": "image_url", "image_url": "https://…/image.png" }`
Both `document_url` and `image_url` accept a public URL or a base64 data URL
(`data:application/pdf;base64,…` / `data:image/png;base64,…`). The `image_url`
field may also be passed as an object: `{ "url": "…" }`.
### Scoping pages [#scoping-pages]
By default the entire document is processed and every page is billed. Use the
optional `pages` field to restrict (and cap the cost of) a request:
* A list of zero-based indices: `"pages": [0, 1, 2]`
* A range string: `"pages": "0-4"`
## curl [#curl]
### Document URL [#document-url]
```bash
curl -X POST "https://api.vichar.io/v1/ocr" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-ocr-latest",
"document": {
"type": "document_url",
"document_url": "https://arxiv.org/pdf/2201.04234"
}
}'
```
### Only specific pages [#only-specific-pages]
```bash
curl -X POST "https://api.vichar.io/v1/ocr" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-ocr-latest",
"document": {
"type": "document_url",
"document_url": "https://arxiv.org/pdf/2201.04234"
},
"pages": "0-4"
}'
```
### Image input [#image-input]
```bash
curl -X POST "https://api.vichar.io/v1/ocr" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-ocr-latest",
"document": {
"type": "image_url",
"image_url": "https://example.com/receipt.png"
}
}'
```
### Inline (base64) document [#inline-base64-document]
```bash
BASE64_PDF=$(base64 -i invoice.pdf)
curl -X POST "https://api.vichar.io/v1/ocr" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"model\": \"mistral-ocr-latest\",
\"document\": {
\"type\": \"document_url\",
\"document_url\": \"data:application/pdf;base64,${BASE64_PDF}\"
}
}"
```
## Response Shape [#response-shape]
```json
{
"pages": [
{
"index": 0,
"markdown": "# Document title\n\nExtracted body text…",
"images": [],
"dimensions": { "dpi": 200, "height": 2200, "width": 1700 }
}
],
"model": "mistral-ocr-latest",
"document_annotation": null,
"usage_info": {
"pages_processed": 1,
"doc_size_bytes": 125344
}
}
```
Each entry in `pages` carries the markdown for one page. `usage_info.pages_processed`
reflects exactly how many pages were billed for the request.
## Billing [#billing]
OCR is billed per page processed, not per token. A request that processes 12
pages bills `12 × $0.004 = $0.048`. Scoping a request with `pages` reduces both
the work and the cost.
## OCR Models Use This Endpoint, Not Chat [#ocr-models-use-this-endpoint-not-chat]
OCR models are not chat models and cannot be called through
`/v1/chat/completions` — doing so returns a `400` pointing you here. Send OCR
requests to `/v1/ocr`.
# Project access
URL: https://docs.vichar.io/features/project-access
Use **Project admin** to let someone manage specific projects without giving them organization administration. Owners and admins assign the role and its project grants from [Team → Members](https://docs.vichar.io/learn/team).
| Role | Scope | Permissions |
| ------------- | ------------------------- | ------------------------------------------------------------------------------------ |
| Owner | All organization projects | Organization administration, billing settings, membership, and project management |
| Admin | All organization projects | Organization and project management; cannot change billing settings or modify owners |
| Project admin | Assigned projects | Project settings, routing, all project API keys, and project-wide usage |
| Developer | Assigned projects | Own API keys and own usage; cannot change project settings |
Project admin and Developer assignments require Enterprise access and at least one project grant. Existing feature and preview requirements still apply. Project admins cannot create or archive projects, change organization settings or provider keys, manage membership, or access organization billing controls. Archiving projects and deleting organizations require an Owner.
## Assign project access [#assign-project-access]
In **Team → Members**, choose **Add Member** or an existing member's **Manage access** action. Select **Project admin**, choose the allowed projects, and save. Invitations carry the selected grants through acceptance. Removing a grant removes access to that project's settings, keys, and usage; the gateway also checks the creator's project access when a key is used, subject to normal cache propagation.
For an authenticated management API session, use `project_admin` with `projectIds` when adding a member:
```json
{
"email": "member@example.com",
"role": "project_admin",
"projectIds": ["your-project-id"]
}
```
Send this body to `POST /team/{organizationId}/members`. Use `PATCH /team/{organizationId}/members/{memberId}` with `role` and the complete `projectIds` list to replace an existing member's access.
## Budgets, teams, and SSO [#budgets-teams-and-sso]
Personal member budgets and IAM rules still apply to the member's keys. Organization teams and default developer budgets apply only to Developers; promoting a Developer clears their team assignment.
Role priority is Owner, Admin, Project admin, then Developer. Project-scoped roles still require explicit project grants: a role mapping does not grant every project.
# Realtime API
URL: https://docs.vichar.io/features/realtime
Vichar supports low-latency, speech-to-speech conversations through the
OpenAI-compatible **`/v1/realtime`** WebSocket endpoint. Sessions support text
and audio input/output, server-side voice activity detection (VAD), input
audio transcription, and function calling — using the same event protocol as
the OpenAI Realtime API. The same endpoint also serves
[transcription-only sessions](#transcription-sessions) for live speech-to-text.
The API is a drop-in replacement: point any OpenAI realtime client at
`wss://api.vichar.io/v1/realtime`, authenticate with your Vichar API
key, and keep your existing event handling.
## Available Models [#available-models]
Browse the available realtime models, with up-to-date pricing, on the
[models page](https://app.vichar.io/dashboard). Realtime sessions are billed per
token — text and audio input, cached input, and output are metered separately
at the model's listed rates, matching the provider's own pricing.
Model names work like everywhere else on Vichar: use the plain model id
(e.g. `gpt-realtime`), a dated alias, or the `provider/model` pinned form
(e.g. `openai/gpt-realtime`).
## Connecting [#connecting]
Connect a WebSocket to:
```
wss://api.vichar.io/v1/realtime?model=gpt-realtime
```
Authenticate with your Vichar API key in the `Authorization` header (an
`x-api-key` header also works):
```javascript
import WebSocket from "ws";
const url = "wss://api.vichar.io/v1/realtime?model=gpt-realtime";
const ws = new WebSocket(url, {
headers: {
Authorization: "Bearer " + process.env.LLM_GATEWAY_API_KEY,
},
});
ws.on("open", () => {
console.log("Connected to server.");
});
ws.on("message", (message) => {
const event = JSON.parse(message.toString());
console.log(event.type);
});
```
Once connected, the session speaks the standard realtime event protocol:
send client events like `session.update`, `conversation.item.create`,
`input_audio_buffer.append`, and `response.create`; receive server events
like `session.created`, `response.output_audio.delta`, and `response.done`.
```javascript
ws.on("open", () => {
// Configure the session.
ws.send(
JSON.stringify({
type: "session.update",
session: {
type: "realtime",
instructions: "You are a friendly assistant.",
audio: {
output: { voice: "marin" },
},
},
}),
);
// Ask for a response.
ws.send(
JSON.stringify({
type: "conversation.item.create",
item: {
type: "message",
role: "user",
content: [{ type: "input_text", text: "Say hello!" }],
},
}),
);
ws.send(JSON.stringify({ type: "response.create" }));
});
```
All events are JSON text frames; audio travels base64-encoded inside events
(binary WebSocket frames are rejected). The model is locked at connection
time — a `session.update` that tries to change `session.model` is rejected
with a `model_locked` error event.
Credentials are never accepted as query parameters. A connection URL
containing `token`, `api_key`, or `client_secret` query parameters is rejected
with HTTP 400. Use the `Authorization` header on servers, or an ephemeral
client secret (below) in browsers.
## Browser Clients and Client Secrets [#browser-clients-and-client-secrets]
Never ship a long-lived API key to a browser. Instead, mint a short-lived
**ephemeral client secret** from your backend with the OpenAI-compatible
`POST /v1/realtime/client_secrets` endpoint:
```bash
curl -X POST "https://api.vichar.io/v1/realtime/client_secrets" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"expires_after": { "anchor": "created_at", "seconds": 120 },
"session": {
"type": "realtime",
"model": "gpt-realtime"
}
}'
```
The response contains the secret (`ek_...`) and its expiry:
```json
{
"value": "ek_...",
"expires_at": 1753500000,
"session": {
"type": "realtime",
"model": "gpt-realtime"
}
}
```
The browser then connects using the standard `openai-insecure-api-key`
WebSocket subprotocol — no `Authorization` header needed:
```javascript
const ws = new WebSocket("wss://api.vichar.io/v1/realtime", [
"realtime",
"openai-insecure-api-key." + clientSecret,
]);
```
Client secret behavior:
* **TTL**: 10–300 seconds (default 60), set via `expires_after.seconds`.
Secrets are reusable until they expire; expiry only gates opening the
connection, not the session's duration.
* **Model pinning**: the secret is minted for one model. The `model` query
parameter is optional when connecting with a secret; if provided, it must
match the minted model.
* **Transcription pinning**: optionally pin an input transcription model at
mint time via `session.audio.input.transcription.model`. The session can
then only enable that transcription model.
* **Instruction pinning**: optionally pin the session prompt at mint time via
`session.instructions` — see below.
* **Voice**: optionally set the session's default output voice at mint time
via `session.audio.output.voice`. Unlike instructions this is only a
default; the client may still change it.
* All authentication, credit, model access, and compliance checks run at
mint time and again at connection time.
### Server-Authoritative Instructions [#server-authoritative-instructions]
By default the session prompt is set by whoever holds the socket, so a browser
client has to send `session.instructions` itself. If your server owns the
prompt, pin it at mint time instead:
```json
{
"session": {
"type": "realtime",
"model": "gpt-realtime",
"instructions": "You are a support agent for ACME."
}
}
```
The gateway then applies the instructions itself when the session connects, so
the prompt never has to travel through the browser. It is not echoed in the
mint response, and it is stripped from the `session.created` and
`session.updated` events the client receives.
Once pinned, the instructions are locked for the session's lifetime, the same
way `session.model` is. A client that tries to change them gets an
`instructions_locked` error event and the change is not applied — this covers
both `session.update` and per-response `response.create.instructions`.
Instructions larger than 16,384 estimated tokens are rejected at mint time with
a 400 `instructions_too_large`, so an oversized prompt fails server-to-server
rather than mid-call.
On Gemini realtime models the same pin applies to `setup.systemInstruction`,
and a client-supplied `systemInstruction` is rejected with the same
`instructions_locked` code.
## Input Audio Transcription [#input-audio-transcription]
Enable transcription of the user's audio with `session.update`. An explicit,
supported transcription model is required — check the
[models page](https://app.vichar.io/dashboard) for available realtime
transcription models and their pricing:
```json
{
"type": "session.update",
"session": {
"type": "realtime",
"audio": {
"input": {
"transcription": { "model": "gpt-4o-transcribe" }
}
}
}
}
```
Transcripts arrive as standard
`conversation.item.input_audio_transcription.delta` and `.completed` events.
* The provider's implicit default (e.g. `whisper-1`) is not available;
omitting `transcription.model` is rejected with
`transcription_model_required`.
* The transcription model is pinned for the session on first use and cannot
be switched afterwards; disabling transcription (`"transcription": null`)
is always allowed.
* Both the current nested form (`session.audio.input.transcription`) and the
legacy top-level form (`session.input_audio_transcription`) are accepted.
* Transcription models are billed the way their provider meters them, per
token or per minute of audio, at the rates on the models page.
## Transcription Sessions [#transcription-sessions]
For live speech-to-text without a speech model in the loop — captions, agent
assist, voice notes — open a **transcription session**. It uses the same
WebSocket endpoint and event protocol, but the session's only model is the
transcription model, so you pay for transcription alone. The model is pinned
at connection time, so it goes in the URL alongside `intent=transcription`:
```
wss://api.vichar.io/v1/realtime?intent=transcription&model=gpt-live-transcribe
```
The gateway applies the transcription model itself once the session is
created. Configure the rest with `session.update`, then stream audio with
`input_audio_buffer.append`:
```json
{
"type": "session.update",
"session": {
"type": "transcription",
"audio": {
"input": {
"format": { "type": "audio/pcm", "rate": 24000 },
"transcription": {
"model": "gpt-live-transcribe",
"delay": "low",
"languages": ["en"]
},
"turn_detection": null
}
}
}
}
```
Transcripts arrive as `conversation.item.input_audio_transcription.delta` and
`.completed` events, exactly as in a realtime session.
* Any realtime transcription model on the
[models page](https://app.vichar.io/dashboard) can open a transcription
session; the page also shows whether a model is billed per token or per
minute of audio.
* The model is locked for the session: a `session.update` that names another
transcription model is rejected with `transcription_model_locked`, and
disabling transcription is rejected with `transcription_model_required`.
Other transcription settings (`delay`, `keywords`, `languages`, `prompt`,
`turn_detection`, `noise_reduction`) are passed through to the provider.
Streaming transcription models segment audio continuously and reject
`turn_detection`: send `null` and commit turns yourself with
`input_audio_buffer.commit`.
* `response.create` is rejected with `response_not_supported`; a
transcription session never generates.
* Browser clients mint a client secret with `session.type: "transcription"`
and connect with the `openai-insecure-api-key` subprotocol as usual. The
`intent` and `model` query parameters are optional with a secret, but if
present they must match what the secret was minted for:
```json
{
"session": {
"type": "transcription",
"audio": {
"input": {
"transcription": { "model": "gpt-live-transcribe" }
}
}
}
}
```
Session gating, limits and billing work as for realtime sessions: each
completed transcription is billed before its transcript is forwarded, and a
session that no longer clears the credit, spend or rate-limit checks is
closed after the current turn. Streaming models transcribe audio before it is
committed; audio that has already produced transcript deltas is committed by
the gateway on disconnect and ahead of an `input_audio_buffer.clear`, so it is
billed like any other turn.
## Function Calling [#function-calling]
Plain function tools work exactly as in the OpenAI Realtime API — declare
them in `session.update` (or per response in `response.create`), receive
`response.function_call_arguments` events, and return results with
`conversation.item.create`:
```json
{
"type": "session.update",
"session": {
"type": "realtime",
"tools": [
{
"type": "function",
"name": "get_weather",
"description": "Get the current weather for a location.",
"parameters": {
"type": "object",
"properties": {
"location": { "type": "string" }
},
"required": ["location"]
}
}
],
"tool_choice": "auto"
}
}
```
Hosted tools (MCP servers, web search, code interpreter, etc.) are not yet
available and are rejected with `tool_type_not_supported`.
## Billing and Session Gating [#billing-and-session-gating]
Realtime sessions bill your organization's pay-as-you-go credits, or your own
provider key when the project uses provider keys. Every
model generation — whether triggered by your `response.create` or by
server-side VAD — passes the gateway's authorization gates first (credits,
API key status, usage limits, IAM rules), so a session cannot run past an
exhausted balance. A blocked generation surfaces as a standard `error` event
on the session instead of a response.
Default per-session safety limits:
| Limit | Default |
| ------------------------------- | ------- |
| Maximum session duration | 1 hour |
| Maximum spend per session | $10 |
| Concurrent sessions per org | 20 |
| Concurrent sessions per API key | 10 |
Sessions that hit a limit are closed gracefully after in-flight responses are
billed. Usage appears in your [activity feed](https://app.vichar.io/dashboard)
like any other request, including separate line items for input transcription.
## Current Limitations [#current-limitations]
* **WebSocket transport only** — WebRTC and SIP are not yet supported.
* **Image input** is not yet supported in realtime sessions.
* **Hosted tools** (MCP, web search, etc.) and **stored prompt references**
(`session.prompt`) are not supported; inline your instructions and function
tools instead.
* Realtime requires a regular developer API key on a pay-as-you-go
organization. End-user session tokens and platform keys are not supported
yet.
## Self-Hosting [#self-hosting]
Realtime is disabled by default on self-hosted deployments. Set
`REALTIME_INLINE=true` to attach the `/v1/realtime` WebSocket listener (and
the client-secret mint endpoint) to the gateway process — `pnpm dev` sets
this automatically. Client secrets require Redis. See `.env.example` for the
tunables (session caps, concurrency limits, shutdown grace period).
# Reasoning
URL: https://docs.vichar.io/features/reasoning
Vichar supports reasoning-capable models that can show their step-by-step thought process before providing a final answer. This feature is particularly useful for complex problem-solving tasks, mathematical calculations, and logical reasoning.
## Reasoning-Enabled Models [#reasoning-enabled-models]
You can find all reasoning-enabled models on our [models page with reasoning filter](https://app.vichar.io/dashboard). These models include:
* OpenAI's GPT-5 series (e.g., `gpt-5`, `gpt-5-mini`)
* Note: GPT-5 models use reasoning but currently do not return the reasoning content in the response.
* Anthropic's Claude 3.7 Sonnet
* Google's Gemini 2.0 Flash Thinking and Gemini 2.5 Pro
* GPT OSS models such as `gpt-oss-120b` and `gpt-oss-20b`
* Z.AI's reasoning models
Some models may reason internally even if the `reasoning_effort` parameter is
not specified.
## Using the Reasoning Parameter [#using-the-reasoning-parameter]
There are two ways to control reasoning effort:
### Option 1: Top-level `reasoning_effort` [#option-1-top-level-reasoning_effort]
Add the `reasoning_effort` parameter directly to your request:
* `none` - Disable reasoning. Supported by OpenAI's newer reasoning models (e.g. `gpt-5.4-mini` and later, which accept `none` instead of `minimal`). For other providers this turns reasoning off.
* `minimal` - Fastest reasoning with minimal thought process (only for GPT-5 models)
* `low` - Light reasoning for simpler tasks
* `medium` - Balanced reasoning for most tasks
* `high` - Deep reasoning for complex problems
* `xhigh` - Very deep reasoning for the most complex problems
* `max` - Highest reasoning tier, above `xhigh`. Supported by Anthropic thinking models and OpenAI GPT-5.6 and later models. Effort tiers are never downgraded by the gateway: providers that accept an effort parameter receive the value unchanged (unsupported values result in a provider error), while providers that take a thinking budget instead (Anthropic, Google, Alibaba) have each tier translated to a native budget
OpenAI's reasoning models do not all accept the same effort values. The
original GPT-5 models support `minimal`, while newer models (e.g.
`gpt-5.4-mini` and later) replace it with `none`. If you send an effort value
the target model doesn't support, OpenAI returns an `unsupported_value` error.
The exact values each provider mapping accepts are exposed as
`reasoning_efforts` on the [`/v1/models`](https://api.vichar.io/v1/models)
endpoint.
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [
{
"role": "user",
"content": "What is 2/3 + 1/4 + 5/6?"
}
],
"reasoning_effort": "medium"
}'
```
### Option 2: Using the `reasoning` object [#option-2-using-the-reasoning-object]
Use the unified `reasoning` configuration object with an `effort` field:
* `none` - Disable reasoning
* `minimal` - Fastest reasoning with minimal thought process
* `low` - Light reasoning for simpler tasks
* `medium` - Balanced reasoning for most tasks
* `high` - Deep reasoning for complex problems
* `xhigh` - Very deep reasoning for the most complex problems
* `max` - Highest reasoning tier, above `xhigh`. Supported by Anthropic thinking models and OpenAI GPT-5.6 and later models. Effort tiers are never downgraded by the gateway: providers that accept an effort parameter receive the value unchanged (unsupported values result in a provider error), while providers that take a thinking budget instead (Anthropic, Google, Alibaba) have each tier translated to a native budget
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5",
"messages": [
{
"role": "user",
"content": "What is 2/3 + 1/4 + 5/6?"
}
],
"reasoning": {
"effort": "medium"
}
}'
```
You cannot use both `reasoning_effort` and `reasoning.effort` in the same
request. Choose one approach. However, you can combine `reasoning_effort` or
`reasoning.effort` with `reasoning.max_tokens` — when `max_tokens` is
specified, it takes priority over the effort level.
### Example Response [#example-response]
The response will include a `reasoning` field in the message object containing the model's step-by-step thought process:
```json
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1234567890,
"model": "gpt-oss-120b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The answer is 1.75 or 7/4.",
"reasoning": "First, I need to find a common denominator for 2/3, 1/4, and 5/6. The LCD is 12. Converting: 2/3 = 8/12, 1/4 = 3/12, 5/6 = 10/12. Adding: 8/12 + 3/12 + 10/12 = 21/12 = 1.75 or 7/4."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"completion_tokens": 45,
"reasoning_tokens": 35,
"total_tokens": 65
}
}
```
## Specifying Reasoning Token Budget [#specifying-reasoning-token-budget]
For models that support it, you can specify an exact token budget for reasoning using the `reasoning` object with `max_tokens`. This gives you precise control over how many tokens the model allocates to its thinking process.
When `reasoning.max_tokens` is specified, it overrides `reasoning.effort` and
`reasoning_effort`. Supported by Anthropic Claude and Google Gemini thinking
models, plus Alibaba-hosted thinking models (forwarded as DashScope's
`thinking_budget`).
### Example Request [#example-request]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4-20250514",
"messages": [
{
"role": "user",
"content": "Explain the P vs NP problem and why it matters."
}
],
"reasoning": {
"max_tokens": 8000
}
}'
```
### Supported Models [#supported-models]
The `reasoning.max_tokens` parameter is supported by:
* **Anthropic Claude**: Claude 3.7 Sonnet, Claude Sonnet 4, Claude Opus 4, Claude Opus 4.5
* **Google Gemini**: Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3 Pro Preview
When using auto-routing or canonical models with `reasoning.max_tokens`, only providers that support this feature will be considered.
### Provider-Specific Constraints [#provider-specific-constraints]
* **Anthropic**: Reasoning budget must be between 1,024 and 128,000 tokens. Values outside this range are automatically clamped.
* **Google**: No specific constraints on the reasoning budget.
### Error Handling [#error-handling]
If you specify `reasoning.max_tokens` for a model that doesn't support it, you'll receive an error:
```json
{
"error": {
"message": "Model gpt-4o does not support reasoning.max_tokens. Remove the reasoning parameter or use a model that supports explicit reasoning token budgets.",
"type": "invalid_request_error",
"code": "model_not_supported"
}
}
```
## Reasoning Mode [#reasoning-mode]
OpenAI's GPT-5.6 and GPT-6 models accept a `reasoning.mode` of `standard` (the default) or `pro`. Pro mode spends additional model work on difficult tasks at higher latency and token usage. It is independent of effort: `mode` selects standard or pro execution, `effort` controls how much reasoning happens within that mode.
```json
{
"model": "gpt-5.6-sol",
"messages": [
{
"role": "user",
"content": "Review this migration plan for failure modes."
}
],
"reasoning": {
"mode": "pro",
"effort": "medium"
}
}
```
The same `reasoning.mode` field works on the `/v1/responses` endpoint. Each mapping lists the values it accepts as `reasoning_modes` on [`/v1/models`](https://api.vichar.io/v1/models); auto-routing only considers mappings that list the requested mode, and a request for a model that does not list it is rejected with a 400 rather than silently run in standard mode:
```json
{
"error": {
"message": "Model gpt-4o does not support reasoning.mode \"pro\". Remove the reasoning.mode parameter or use a model whose reasoning_modes on /v1/models include it.",
"type": "invalid_request_error",
"code": "model_not_supported"
}
}
```
## Controlling Response Verbosity [#controlling-response-verbosity]
For OpenAI GPT-5 and later models, you can control how detailed the model's final answer is with the top-level `verbosity` parameter. This is independent of `reasoning_effort`: `reasoning_effort` controls how much the model thinks, while `verbosity` controls how much it writes in its response.
Accepted values:
* `low` - Concise responses with minimal elaboration
* `medium` - Balanced level of detail
* `high` - Detailed, thorough responses
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5",
"messages": [
{
"role": "user",
"content": "Explain how a hash map works."
}
],
"verbosity": "low"
}'
```
`verbosity` is only supported by OpenAI GPT-5 and later models. You can check
which model mappings accept it via the
[`/v1/models`](https://api.vichar.io/v1/models) endpoint. It can be combined
freely with `reasoning_effort` or the `reasoning` object.
### Error Handling [#error-handling-1]
If you specify `verbosity` for a model that doesn't support it, you'll receive a `400` error:
```json
{
"error": {
"message": "Model gpt-4o does not support the verbosity parameter. Remove the verbosity parameter or use a model that supports it (OpenAI GPT-5 and later).",
"type": "invalid_request_error",
"code": "model_not_supported"
}
}
```
## Streaming Reasoning Content [#streaming-reasoning-content]
When streaming is enabled, reasoning content will be streamed as part of the response chunks:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [
{
"role": "user",
"content": "Solve this logic puzzle: If all roses are flowers and some flowers fade quickly, can we conclude that some roses fade quickly?"
}
],
"reasoning_effort": "high",
"stream": true
}'
```
The reasoning content will appear in the stream chunks before the final answer, allowing you to display the model's thought process in real-time.
Example:
```
data: {
"id": "chatcmpl-fb266880-1016-4797-9a70-f21a538edaf6",
"object": "chat.completion.chunk",
"created": 1761048126,
"model": "openai/gpt-oss-20b",
"choices": [
{
"index": 0,
"delta": {
"reasoning": "It's ",
"role": "assistant"
},
"finish_reason": null
}
]
}
```
## Preserving Reasoning Across Calls [#preserving-reasoning-across-calls]
Agents work over many turns: the model reasons, calls a tool, and the tool result arrives in the next request. Reasoning models are trained to continue from the reasoning they produced on earlier turns, so dropping it mid-task degrades tool-calling and multi-turn performance.
Vichar preserves reasoning the same way OpenAI's Responses API does. The model returns an opaque **encrypted reasoning** payload, and you send it back unchanged along with the rest of the conversation on the next call. The payload is issued and verified by the provider — the gateway passes it through without inspecting or rewriting it.
Readable reasoning summaries and opaque replay data are separate. Summaries
remain available when the model provides them; the gateway cannot decrypt
provider-issued reasoning payloads or Gemini thought signatures.
### Responses API [#responses-api]
Ask for the payloads with `include: ["reasoning.encrypted_content"]`, and send `store: false` if you want the turn to stay stateless:
```bash
curl -X POST "https://api.vichar.io/v1/responses" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.4-nano",
"input": [{ "role": "user", "content": "What is the weather in Paris? Use the tool." }],
"tools": [
{
"type": "function",
"name": "get_weather",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
],
"store": false,
"include": ["reasoning.encrypted_content"],
"reasoning": { "effort": "high" }
}'
```
Reasoning items in `output` then carry `encrypted_content`:
```json
{
"type": "reasoning",
"id": "rs_057d93d1ba22357e01...",
"summary": [],
"encrypted_content": "gAAAAABqfkEO3uEFtXeGSNE9..."
}
```
On the next turn, replay the previous `output` items unchanged — reasoning items included — followed by your tool results:
```json
{
"model": "gpt-5.4-nano",
"input": [
{
"role": "user",
"content": "What is the weather in Paris? Use the tool."
},
{
"type": "reasoning",
"id": "rs_057d93d1ba22357e01...",
"summary": [],
"encrypted_content": "gAAAAABqfkEO3uEFtXeGSNE9..."
},
{
"type": "function_call",
"call_id": "call_923JIrDbCvIHnGqfvDbPeaLh",
"name": "get_weather",
"arguments": "{\"city\":\"Paris\"}"
},
{
"type": "function_call_output",
"call_id": "call_923JIrDbCvIHnGqfvDbPeaLh",
"output": "18C sunny"
}
],
"store": false,
"include": ["reasoning.encrypted_content"],
"reasoning": { "effort": "high" }
}
```
Without the `include` value, reasoning items come back without `encrypted_content` and there is nothing to replay. When streaming, the payload arrives on the reasoning item's `response.output_item.done` event.
#### Stateful Alternative [#stateful-alternative]
If you would rather not carry the payloads yourself, leave `store` at its default (`true`) and chain turns with `previous_response_id`. The gateway keeps the stored response — encrypted reasoning included — for [30 days](https://docs.vichar.io/features/data-retention#retention-periods) and replays the reasoning for you, so `include` is not needed on that path.
### Chat Completions [#chat-completions]
The Chat Completions format has no reasoning items, so the gateway carries the same payloads on the assistant message as `reasoning_details`:
```json
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_923JIrDbCvIHnGqfvDbPeaLh",
"type": "function",
"function": { "name": "get_weather", "arguments": "{\"city\":\"Paris\"}" }
}
],
"reasoning_details": [
{
"type": "reasoning.encrypted",
"data": "gAAAAABqfkEO3uEFtXeGSNE9...",
"id": "rs_057d93d1ba22357e01...",
"format": "openai-responses-v1",
"index": 0
}
]
}
```
To preserve the reasoning, append the assistant message back into `messages` exactly as you received it (with `reasoning_details` intact) before the matching `tool` result. No opt-in parameter is required — the field is always present when the model produced one. When streaming, the entries arrive as a `delta.reasoning_details` chunk.
The gateway converts supported `reasoning_details` formats to native provider fields and strips these entries from other upstream requests.
### Gemini Thought Signatures [#gemini-thought-signatures]
Gemini text turns carry signatures in `reasoning_details` entries with `type: "reasoning.text"` and `format: "google-gemini-v1"`. Replay the entire assistant message, including these entries. Preserve each entry's `google_part` metadata: it lets the gateway restore the original signed text boundaries. Streaming clients must collect `delta.reasoning_details`, including entries arriving with empty text.
When the gateway repairs JSON, clients receive the repaired answer. The entries also carry `google_response` metadata so replay can restore the original signed text. Editing the returned answer prevents that restoration.
On the Responses API, replay the complete `output`: assistant message items carry `reasoning_details`, while function call items carry `extra_content.google.thought_signature`. These fields are returned without an `include` opt-in. Stored conversations using `previous_response_id` preserve them automatically. Gemini signatures are not converted into OpenAI `encrypted_content`.
For Chat Completions, `content[]` text parts and `tool_calls[]` also accept `extra_content.google.thought_signature`. Streaming tool calls have distinct `index` values across chunks; accumulate calls by index and keep each signature with its original call. The tool-call signature cache remains a fallback, but replaying the metadata also works when signatures are not cached. Only replay signatures with the original content and the model/provider that issued them.
### Error Handling [#error-handling-2]
Payloads are verified by the provider that issued them. A modified, truncated, or foreign payload is rejected upstream and the error is forwarded unchanged:
```json
{
"error": {
"message": "The encrypted content for item rs_057d93d1ba22357e01... could not be verified. Reason: Encrypted content could not be decrypted or parsed.",
"type": "invalid_request_error",
"code": "invalid_encrypted_content"
}
}
```
Only replay payloads against the model and provider that produced them.
## Usage Tracking [#usage-tracking]
### Response Payload [#response-payload]
The `usage` object in the response includes reasoning-specific token counts:
* `reasoning_tokens` - Number of tokens used for the reasoning process
* `completion_tokens` - Number of tokens in the final answer
* `prompt_tokens` - Number of tokens in the input
* `total_tokens` - Sum of all token counts
### Logs and Analytics [#logs-and-analytics]
All requests using the `reasoning_effort` parameter are tracked in your dashboard logs with:
* The `reasoningContent` field containing the full reasoning text
* Separate token counts for reasoning vs. completion
* Performance metrics for reasoning-enabled requests
You can view detailed logs for each request in the [dashboard](https://app.vichar.io/dashboard) to analyze how models are reasoning through problems.
## Auto-Routing with Reasoning [#auto-routing-with-reasoning]
When using auto routing (`"model": "auto"`), if the selected model supports reasoning, you did not set a reasoning effort yourself, and the request has no web search tool (web search is incompatible with `minimal` effort), Vichar will:
1. Automatically set `reasoning_effort` to `minimal` for GPT-5 models
2. Set `reasoning_effort` to `low` for other auto-routed reasoning models
3. Only route to providers that support reasoning when `reasoning_effort` is specified
This ensures optimal performance and cost when using auto-routing with reasoning-capable models.
## Model-Specific Behavior [#model-specific-behavior]
Not all reasoning models return reasoning content in the same way. Some models (like OpenAI models) may reason internally but not expose the reasoning content in the response. Vichar makes sure the response is unified across different providers, but the depth and format of reasoning may vary.
## Best Practices [#best-practices]
1. **Choose appropriate reasoning effort**: Use `low` or `minimal` for simple tasks, `medium` for most tasks, and `high` only for complex problems that require deep reasoning
2. **Monitor token usage**: Reasoning can significantly increase token consumption - monitor your `reasoning_tokens` in the usage object
3. **Stream for better UX**: When building user-facing applications, enable streaming to show the reasoning process in real-time
4. **Check logs**: Review the `reasoningContent` in your dashboard logs to understand how models are solving problems
## Error Handling [#error-handling-3]
If you specify `reasoning_effort` for a model that doesn't support reasoning, you'll receive an error:
```json
{
"error": {
"message": "Model gpt-4o does not support reasoning. Remove the reasoning_effort parameter or use a reasoning-capable model.",
"type": "invalid_request_error",
"code": "model_not_supported"
}
}
```
To avoid this error, only use the `reasoning_effort` parameter with [reasoning-enabled models](https://app.vichar.io/dashboard).
# Rerank
URL: https://docs.vichar.io/features/rerank
Vichar exposes a Cohere-compatible `/v1/rerank` endpoint that scores a list
of candidate documents against a query and returns them ordered by relevance.
Rerankers are cross-encoders: they read the query and a document *together*
rather than embedding each in isolation. That makes them far more accurate than
vector similarity, but too slow to run over a whole corpus. The usual pattern is
two-stage retrieval — use [embeddings](https://docs.vichar.io/features/embeddings) to cheaply fetch
the top \~100 candidates, then rerank those down to the handful you actually put
in the prompt.
Browse available rerank models on the
[models page](https://app.vichar.io/dashboard).
For the full request and response schema, see the
[API reference](https://docs.vichar.io/v1_rerank).
## Endpoint [#endpoint]
`POST https://api.vichar.io/v1/rerank`
## cURL [#curl]
```bash
curl -X POST "https://api.vichar.io/v1/rerank" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-reranker-8b",
"query": "What is the capital of France?",
"documents": [
"Paris is the capital of France.",
"Berlin is the capital of Germany.",
"Madrid is the capital of Spain."
],
"top_n": 2
}'
```
```json
{
"id": "kK1sT9pQ2mR7vX4nB6cL8dF3gH5jW0yZaA2eS4uI",
"results": [
{ "index": 0, "relevance_score": 0.9781517386436462 },
{ "index": 1, "relevance_score": 0.00010548105638008565 }
],
"meta": {
"api_version": { "version": "1" },
"billed_units": { "input_tokens": 255, "total_tokens": 255 }
}
}
```
Results are sorted by `relevance_score` descending. Each `index` refers back to
the position of that document in the request's `documents` array, so you can map
scores onto your own records without relying on the returned text.
The response `id` echoes the `x-request-id` header when you send one, which is
handy for correlating a rerank call with its entry in the activity log.
## Request fields [#request-fields]
| Field | Type | Description |
| -------------------- | ---------- | ------------------------------------------------------------------------- |
| `model` | `string` | Rerank model to use. Optionally prefixed with a provider (`deepinfra/…`). |
| `query` | `string` | The search query to rank documents against. |
| `documents` | `string[]` | Candidate documents to score. At least one. |
| `top_n` | `number` | Return only the top N results. Defaults to all documents. |
| `return_documents` | `boolean` | Include the document text on each result. Defaults to `false`. |
| `max_chunks_per_doc` | `number` | Maximum chunks per document. Accepted, but not every provider honors it. |
## Two-stage retrieval [#two-stage-retrieval]
```ts
// 1. Cheap recall: embed the query and pull the top 100 candidates
// from your vector store.
const candidates = await vectorStore.search(queryEmbedding, { limit: 100 });
// 2. Precise ordering: rerank those candidates and keep the best 5.
const res = await fetch("https://api.vichar.io/v1/rerank", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "qwen3-reranker-8b",
query,
documents: candidates.map((c) => c.text),
top_n: 5,
}),
});
const { results } = await res.json();
const topDocuments = results.map((r) => candidates[r.index]);
```
Rerank models are billed on input tokens only — the query and every document
you send are scored as input. There are no output tokens, since the response
is a list of scores rather than generated text. Sending 100 long documents
costs real money, so keep the first-stage candidate set tight.
Rerank models only work on `/v1/rerank`. Requesting one on
`/v1/chat/completions` returns a 400 pointing you at the right endpoint, and
they are not available in the playground.
# Response Healing
URL: https://docs.vichar.io/features/response-healing
Response Healing is a plugin that automatically validates and repairs malformed JSON responses from AI models. When enabled, Vichar ensures that API responses conform to your specified schemas even when the model's formatting is imperfect.
## Why Response Healing? [#why-response-healing]
Large language models occasionally produce invalid JSON, especially in complex scenarios:
* **Markdown wrapping**: Models often wrap JSON in code blocks like \`\`\`json...\`\`\`
* **Mixed content**: JSON may be preceded or followed by explanatory text
* **Syntax errors**: Trailing commas, unquoted keys, or single quotes instead of double quotes
* **Truncated output**: Token limits may cut off responses mid-JSON
Response Healing automatically detects and fixes these issues, saving you from implementing error handling for every possible malformed response.
## Enabling Response Healing [#enabling-response-healing]
To enable Response Healing, add `response-healing` to the `plugins` array in your request:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Return a JSON object with name and age"}],
"response_format": {"type": "json_object"},
"plugins": [{"id": "response-healing"}]
}'
```
Response Healing only activates when `response_format` is set to `json_object`
or `json_schema`. For regular text responses, the plugin has no effect.
## How It Works [#how-it-works]
When Response Healing is enabled, Vichar applies a series of repair strategies to malformed JSON responses:
### 1. Markdown Extraction [#1-markdown-extraction]
Extracts JSON from markdown code blocks:
```text
Here's the data:
\`\`\`json
{"name": "Alice", "age": 30}
\`\`\`
```
Becomes:
```json
{ "name": "Alice", "age": 30 }
```
### 2. Mixed Content Extraction [#2-mixed-content-extraction]
Separates JSON from surrounding text:
```text
Sure! Here is the JSON you requested: {"name": "Alice", "age": 30} Let me know if you need anything else.
```
Becomes:
```json
{ "name": "Alice", "age": 30 }
```
### 3. Syntax Fixes [#3-syntax-fixes]
Repairs common JSON syntax violations:
| Issue | Before | After |
| --------------- | ------------------- | ------------------- |
| Trailing commas | `{"a": 1,}` | `{"a": 1}` |
| Unquoted keys | `{name: "Alice"}` | `{"name": "Alice"}` |
| Single quotes | `{'name': 'Alice'}` | `{"name": "Alice"}` |
### 4. Truncation Completion [#4-truncation-completion]
Adds missing closing brackets for truncated responses:
```text
{"name": "Alice", "data": {"nested": true
```
Becomes:
```json
{ "name": "Alice", "data": { "nested": true } }
```
## Usage Examples [#usage-examples]
### With JSON Object Format [#with-json-object-format]
Request a structured response with automatic healing:
```typescript
const response = await fetch("https://api.vichar.io/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-4o",
messages: [
{
role: "user",
content:
"Return a JSON object with fields: name (string) and age (number)",
},
],
response_format: { type: "json_object" },
plugins: [{ id: "response-healing" }],
}),
});
const result = await response.json();
// Response is guaranteed to be valid JSON
const data = JSON.parse(result.choices[0].message.content);
```
### With JSON Schema [#with-json-schema]
For stricter validation, combine with `json_schema`:
```typescript
const response = await fetch("https://api.vichar.io/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-4o",
messages: [
{
role: "user",
content: "Generate a user profile",
},
],
response_format: {
type: "json_schema",
json_schema: {
name: "user_profile",
schema: {
type: "object",
required: ["name", "email"],
properties: {
name: { type: "string" },
email: { type: "string" },
age: { type: "number" },
},
},
},
},
plugins: [{ id: "response-healing" }],
}),
});
const result = await response.json();
```
## Healing Metadata [#healing-metadata]
When a response is healed, the healing method is logged for debugging. The following healing methods may be applied:
| Method | Description |
| -------------------------- | ------------------------------------------- |
| `markdown_extraction` | JSON extracted from markdown code blocks |
| `mixed_content_extraction` | JSON extracted from surrounding text |
| `syntax_fix` | Trailing commas, quotes, or keys were fixed |
| `truncation_completion` | Missing closing brackets were added |
| `combined_strategies` | Multiple strategies were applied |
## Limitations [#limitations]
Response Healing works for streaming requests too, but it changes how the
stream is delivered: the gateway buffers the whole content stream, repairs it,
and replays the healed content as a single chunk — so token-by-token streaming
is lost for that request. Healing is disabled for multi-choice streams (`n`
greater than 1).
Healing also runs automatically — without the `plugins` entry — for a few
providers and models whose native JSON mode is known to emit malformed output;
in those cases the gateway repairs `response_format` JSON responses on its own.
Response Healing works best for:
* Simple to moderately complex JSON structures
* Common formatting issues from LLMs
It may not be able to repair:
* Severely corrupted or nonsensical output
* Complex nested structures with multiple issues
* Responses that don't contain any recognizable JSON
## Best Practices [#best-practices]
### Use with Structured Prompts [#use-with-structured-prompts]
Combine Response Healing with clear instructions for best results:
```typescript
const response = await fetch("https://api.vichar.io/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-4o",
messages: [
{
role: "system",
content: "Always respond with valid JSON. No explanations.",
},
{
role: "user",
content: "List three colors as a JSON array",
},
],
response_format: { type: "json_object" },
plugins: [{ id: "response-healing" }],
}),
});
const result = await response.json();
```
### Validate Critical Data [#validate-critical-data]
For critical applications, validate the healed JSON in your code:
```typescript
const result = await response.json();
const content = result.choices[0].message.content;
const data = JSON.parse(content);
// Add your own validation
if (!data.name || typeof data.name !== "string") {
throw new Error("Invalid response: missing name");
}
```
### Monitor Healing Rates [#monitor-healing-rates]
If you notice frequent healing in your logs, consider:
* Improving your prompts to request cleaner JSON
* Using models with better JSON output (e.g., GPT-4o, Claude 3.5)
* Adding explicit JSON examples in your prompts
# Routing
URL: https://docs.vichar.io/features/routing
Vichar provides flexible and intelligent routing options to help you get the best performance and cost efficiency from your AI applications. Whether you want to use specific models, providers, or let our system automatically optimize your requests, we've got you covered.
Vichar also includes **automatic retry and fallback** — if a provider fails, your request is seamlessly retried on the next best provider, all within the same API call.
## Global Provider Rate Limits [#global-provider-rate-limits]
Administrators can set a global RPM or RPD limit to `0` to block matching
provider/model requests, with either **Global** or **Per-organization**
enforcement. Remove the limit or set a positive value to resume traffic; changes
propagate through the rate-limit cache (default 60 seconds).
Existing precedence still applies: organization-specific limits and more specific
global provider/model limits can override a provider-wide limit for the same
window. Routing can fall back to another available provider. If a zero-limited
provider remains selected, the request returns `429` without calling it, even
when every candidate is rate-limited. No configured limit still means unlimited;
other rate-limit settings retain their existing behavior.
## Model Selection [#model-selection]
### Any Model Name [#any-model-name]
You can use any model name from our [models page](https://app.vichar.io/dashboard) or discover available models programmatically through the [/v1/models endpoint](https://docs.vichar.io/v1_models).
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
### Model ID Routing [#model-id-routing]
Choose a specific model ID to route to the **best available provider** for that model. Vichar's smart routing algorithm considers multiple factors to find the optimal provider across all configured options.
#### Smart Routing Algorithm [#smart-routing-algorithm]
When you use a model ID without a provider prefix, Vichar's intelligent routing system analyzes multiple factors to select the best provider.
**Weighted Scoring System**:
Each factor has a **relative weight**. The factors are scored as ratios against the best provider in the candidate set (e.g. a provider that is twice as expensive as the cheapest scores `1.0` on price), and each ratio is multiplied by its weight divided by the sum of all active weights. The provider with the lowest (best) total score wins.
The default weights are:
| Factor | Default weight | Notes |
| --------------- | -------------- | --------------------------------------------------------------------------------- |
| **Price** | `0.6` | Token cost for the expected input/output mix, including cache reads when relevant |
| **Uptime** | `0.5` | Provider reliability / low error rate |
| **Throughput** | `0.05` | Tokens per second generation speed |
| **Latency** | `0.025` | Time to first token — **only applied for streaming requests** |
| **Cache** | `0` | Optional cache-support preference; cache-read savings already count toward price |
| **Image price** | `1.0` | Replaces the price weight for image-generation models |
Because the weights are relative and normalized by the sum of the active weights, price and uptime dominate routing decisions in practice, while throughput and latency act as tie-breakers between otherwise comparable providers.
**Latency Weight for Non-Streaming Requests**:
The latency weight only applies to streaming requests (time-to-first-token is only measured there). For non-streaming requests the latency weight is dropped and its share is redistributed proportionally across the remaining factors.
**Time-Decayed Metrics Window**:
Provider metrics (uptime, throughput, latency) are not a flat "last N minutes" snapshot. They are aggregated over a rolling **60-minute window** with a time-decay weighting so very recent behavior dominates while older data still contributes:
* The most recent **1 minute** is weighted **10×**
* The most recent **5 minutes** are weighted **3×**
* The remainder of the 60-minute window is weighted **1×**
This makes routing react quickly to a provider that just started failing or slowing down, without overreacting to a single noisy data point.
**Prompt Caching and Token Costs**:
For estimated prompts of at least **5,000 tokens**, or when choosing a session's provider, routing blends each provider's uncached and cached input prices and weights output by the expected output:input token ratio. Coding sessions that mostly reuse prompt tokens can therefore favor a provider with cheaper cache reads even when its uncached input price is higher. Providers without a cached input price use their full input price.
These estimates use the project's **last 24 hours of model usage**, once it includes at least **20 successful requests and 20,000 input tokens**:
* **Cache-hit rate** is cached input tokens divided by total input tokens. Each provider uses its own rate, counted across all of its regions, once it meets the same sample thresholds; otherwise it uses the project's combined rate for that model.
* **Output:input ratio** uses the project's combined output and input tokens for that model across providers.
Routing caches the usage lookup for 60 seconds. It reads hourly aggregates, which work with payload retention disabled. Hourly buckets containing gateway response-cache hits are excluded because they cannot isolate upstream usage, so projects with response caching enabled may keep using the defaults. During lookup failures, routing uses previously cached observations when available, then falls back to configured estimates. These are workload estimates, not guarantees that a particular prompt will hit a provider's cache.
Without enough history, routing uses these initial workload estimates:
| Workload | Cached input | Output:input ratio |
| -------------------------- | ------------ | ------------------ |
| General API / unknown | 10% | 20% |
| A recognized coding client | 90% | 2% |
| Chat organization | 50% | 10% |
Recognized coding clients use the coding defaults even on regular API projects. A session id alone does not identify coding traffic. These are starting assumptions; sufficient project/model observations replace them. The dashboard reports organization defaults, while recognized coding requests use the coding profile at request time.
Explicit Enterprise overrides for `thresholds.cacheHitRate` and `thresholds.cacheOutputRatio` take precedence over both workload defaults and observations. Setting them to `0` and `1`, respectively, restores list-price ranking.
Both `auto` and `price` routing use these token-cost estimates. The separate **cache weight** defaults to `0`, so cache support alone does not outweigh lower estimated costs. Enterprise projects can explicitly enable that additional preference under `auto`; `price` routing always sets it to zero. Cache support appears as `cacheSupported` in routing metadata.
When choosing a [session's provider](#sticky-session-routing), routing applies the workload estimate even to a short opening prompt, using observations when available and workload defaults otherwise. This estimates the session's token mix; the opening request may still incur cache misses. Small requests outside a session weight input and output prices equally and omit the cache weight.
**Exponential Uptime Penalty**:
Providers with uptime below 95% receive an additional exponential penalty that increases rapidly as uptime drops:
* 95-100% uptime: No penalty
* 90% uptime: \~0.07 penalty
* 80% uptime: \~0.62 penalty
* 70% uptime: \~1.73 penalty
* 50% uptime: \~5.61 penalty
This ensures providers experiencing significant issues are strongly deprioritized while minor fluctuations have minimal impact. The penalty threshold (default `95%`) is configurable.
**Provider Priority**:
Each provider has a **priority** value (default `1`) that nudges routing toward or away from it independently of live metrics:
* A provider's priority is applied as a `(1 - priority)` adjustment to its score — higher priority lowers the score (more preferred), lower priority raises it (less preferred).
* A priority of **0** disables the provider entirely, removing it from routing for that model.
Provider priorities are surfaced in the routing metadata so you can see how they influenced a decision.
**Epsilon-Greedy Exploration** (1% of requests by default):
To solve the "cold start problem" where new or unused providers never get traffic to build up metrics, the system randomly explores different providers a small fraction of the time (default 1%, configurable). This ensures:
* All providers periodically receive traffic
* New providers can prove their reliability
* The system adapts to changing provider performance
* You benefit from improved routing decisions over time
The exploration rate is configurable per project through the routing configuration (`thresholds.explorationRate`), and self-hosted deployments can override it globally with the `EXPLORATION_RATE` environment variable (a number between `0` and `1`).
**Stable Provider Preference**:
To avoid unnecessary churn between providers that score similarly, Vichar remembers the best provider chosen for each model and sticks with it across requests — even if another provider edges ahead slightly on the next score calculation.
On every routing decision, the system checks whether the previously selected provider is still acceptable:
* **Uptime hard switch**: if the preferred provider's uptime drops below **85%**, routing switches to the current best-scoring provider immediately.
* **Score margin soft switch**: the preferred provider is replaced only when a better option's score is more than **0.15** ahead. Small fluctuations caused by metric noise or minor price differences do not trigger a switch.
* **Periodic re-evaluation**: the preference expires after **1 hour**, at which point the next request picks the best-scoring provider fresh and stores it as the new preferred.
Requests that are part of the epsilon-greedy exploration bypass this preference entirely so that all providers continue to receive periodic traffic and build up metrics.
The selection reason in routing metadata will show `stable-preferred` when a request was served by the stored preference rather than the top-scored provider at that moment.
Self-hosted deployments can tune this behavior with three environment
variables: `PREFERRED_PROVIDER_TTL` (preference lifetime in seconds, default
`3600`), `PREFERRED_PROVIDER_UPTIME_THRESHOLD` (hard-switch uptime floor,
default `85`), and `PREFERRED_PROVIDER_SCORE_MARGIN` (soft-switch score gap,
default `0.15`). On the **Enterprise plan**, these same values can be
customized per project from the dashboard — see [Per-Project Routing
Configuration](#per-project-routing-configuration-enterprise).
**Routing Metadata**:
Every request includes detailed routing metadata in the logs, showing:
* Available providers that were considered
* Selected provider and selection reason
* Scores for each provider (including uptime, throughput, latency, price, priority, and cache support)
This transparency allows you to understand and debug routing decisions.
Using model IDs without a provider prefix automatically routes to the optimal
provider based on reliability, speed, and cost. The system continuously learns
and adapts based on real-time performance metrics.
Smart routing prioritizes reliability over cost, ensuring your requests are
routed to providers with proven uptime and performance, while still
considering cost efficiency.
### Routing Strategy [#routing-strategy]
By default, model-ID routing uses the full weighted score described above (`routing: "auto"`). When you care about a single dimension, set the `routing` field — named after the factor it optimizes — to bias provider selection toward it:
| Strategy | Behavior |
| ---------------------------- | ------------------------------------------------------------------------------------------ |
| `auto` *(default)* | Full weighted smart-routing score (price, uptime, throughput, latency, cache). |
| `price` | Gives price a **90% relative weight**, including estimated cache-read costs when relevant. |
| `throughput` | Gives throughput a **90% relative weight**, so the fastest-generating provider wins. |
| `latency` | Gives latency a **90% relative weight**, so the lowest time-to-first-token wins. |
Each non-`auto` strategy keeps a small (10%) uptime weight, and the [exponential uptime penalty](#smart-routing-algorithm) still applies on top. This means the dominant pick is still skipped in favor of another provider when it has extremely bad uptime — you get the cheapest (or fastest) provider that is actually healthy, not one that is effectively down.
Because time-to-first-token is only measured for streaming requests, `routing: "latency"` only biases streaming requests; for non-streaming requests it falls back to selecting on uptime.
```bash
# Always pick the cheapest healthy provider for this model
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}],
"routing": "price"
}'
```
```bash
# Always pick the highest-throughput healthy provider for this model
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}],
"routing": "throughput"
}'
```
The `routing` field only applies to model-id routing. Combining it with a
specific provider (e.g. `openai/gpt-4o`) returns a `400` error, since the
strategy can't influence a pinned provider — remove the provider prefix to use
a strategy. On **coding (dev) plans**, only `auto` and `price` are allowed;
the other strategies return a `400` error because they would bypass the
prompt-cache–aware routing those plans depend on.
### Sticky Session Routing [#sticky-session-routing]
When a model is served by multiple providers, every request is normally scored independently — so a multi-turn conversation can bounce between providers. That defeats provider-side **prompt caching**, which only pays off when consecutive requests with a shared prefix hit the **same** provider.
Sticky session routing solves this: attach a session identifier and Vichar pins all requests for that session to a single provider (and region), keeping the upstream prompt cache warm across the whole conversation.
#### Setting the session id [#setting-the-session-id]
For chat completions, the session key is resolved in priority order:
1. The `x-session-id` header
2. The `x-session-affinity` header (sent automatically by coding agents such as opencode)
3. The `session_id` or `session-id` header
4. The `prompt_cache_key` body field (OpenAI-compatible)
5. The `user` body field (OpenAI-compatible)
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-H "x-session-id: conversation-9f8e7d6c" \
-d '{
"model": "claude-sonnet-4-6",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
For the Anthropic Messages endpoint (`/v1/messages`), the session key is derived automatically from `metadata.user_id` — coding agents such as Claude Code embed the session id there — and forwarded internally. An explicit `x-session-id` header still takes precedence.
#### How pinning works [#how-pinning-works]
On a session's **first** request the provider is chosen by the normal weighted smart-routing score — the same price-, priority-, uptime-, and throughput-aware algorithm used for non-sticky requests. That choice is then **persisted for the session** and reused on every subsequent request, so the upstream prompt cache stays warm without bouncing the conversation between providers.
The first selection uses the expected cache-hit rate and output/input mix even if the opening prompt is short. Routing uses observed project/model usage when sufficiently sampled, otherwise the workload defaults above. New usage observations affect future provider selections; a healthy existing pin stays in place as described below.
Because the pinned provider is replayed directly, sticky requests **skip the epsilon-greedy exploration** — a session is never randomly bounced to a different provider mid-conversation.
Request compatibility takes precedence over the saved pin. The gateway first filters mappings for requirements such as input modalities, service tiers, regions, and a non-`auto` `tool_choice`, then looks for the pinned provider in that eligible set. If the pinned mapping cannot honor the request but another mapping can, the session moves to the capable mapping and the pin is updated. For a fixed model or dynamic route where **no** mapping can honor `tool_choice`, the gateway preserves availability instead: it keeps the mappings, downgrades `tool_choice` to `auto`, and sticky routing may retain the existing pin. Automatic model selection does not use that fallback because it can choose a capable model instead.
#### Falling back when a provider is down [#falling-back-when-a-provider-is-down]
An established pin yields only when its provider can no longer serve the session well. A session is re-scored and re-pinned to the current weighted-best provider when its provider:
* Drops below the session uptime threshold (default 85%), except for Gemini sessions whose signatures require provider affinity, or
* Is filtered out of the candidate set (health or compatibility filtering).
Sticky requests never enter the cross-provider [automatic retry & fallback](#automatic-retry--fallback) loop — a transient failure is retried against the pinned provider only, on another configured key when several exist, or on the same platform key when only one is configured. The failure still degrades that provider's uptime metrics, which is what triggers re-pinning on a subsequent request once the uptime threshold is crossed.
Re-pinning runs the same weighted algorithm again, so the replacement is the best currently available provider — not an arbitrary one.
Gemini sessions keep their eligible provider through uptime dips because thought signatures cannot be replayed across Google API providers. If a pin expires or the provider becomes ineligible, moving a conversation with its old signatures can produce `Corrupted thought signature` or `Invalid thought signature`. On the pay-as-you-go API, keep a Gemini conversation on one provider with a provider-prefixed model and `X-No-Fallback: true` from its first request. For an affected conversation, resume with the provider that issued its signatures or start a new conversation. See [reasoning replay](https://docs.vichar.io/features/reasoning#gemini-thought-signatures).
The selection reason in routing metadata shows `session-sticky` when a request was pinned via a session id.
Sticky routing optimizes for cache locality over per-request churn. Once a
session is pinned it stays on its provider even if a cheaper or faster
alternative becomes momentarily available, since the prompt-cache savings
typically outweigh the difference — but the initial pick still respects price
and priority. Requests without a session id are unaffected and continue to use
the weighted smart-routing algorithm.
### Provider-Specific Routing [#provider-specific-routing]
To use a specific provider without any fallbacks, prefix the model name with the provider name followed by a slash:
```bash
# Use OpenAI specifically
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'
# Use DeepSeek provider specifically
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
#### Regions [#regions]
Some providers expose the same model in multiple regions. In that case, Vichar supports two routing modes:
* `provider/model` selects the best eligible region for that provider using the same routing inputs used elsewhere: recent uptime, throughput, latency, and price
* `provider/model:region` pins the request to one exact region
```bash
# Let Vichar choose the best Alibaba region for DeepSeek V3.2
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/deepseek-v3.2",
"messages": [{"role": "user", "content": "Hello!"}]
}'
# Force a specific Alibaba region
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "alibaba/deepseek-v3.2:cn-beijing",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
If your provider key stores an explicit region, that region acts like a lock and Vichar will only use that region for provider-specific requests. If no explicit region is configured on the provider key, provider-specific requests can still score all eligible regions for that provider.
Routing metadata reflects this:
* Dynamic provider-region selection shows all eligible regional scores that were considered
* Explicitly pinned regions show only the pinned region in the score list
Region-aware routing only compares regions that are actually available for the
current project mode and provider setup. In credits mode, that means only
regions backed by configured environment keys. In API keys and hybrid mode, an
explicit provider-key region restricts the request to that region.
A few regions are served by an endpoint that belongs to your own account rather
than a shared one — Alibaba Cloud's EU (Frankfurt) region has no shared
DashScope domain and is served by your Model Studio workspace's dedicated host.
Such a region still works from an API key alone, via the provider's shared entry
point, but that endpoint is rate-limited and carries no SLA. Set the workspace
ID on the provider key (copy it from the API Host shown when you create the key)
to route through your own endpoint instead.
#### Low-Uptime Protection [#low-uptime-protection]
When you specify a provider explicitly, Vichar checks the provider's recent uptime (from the time-decayed metrics window described above). If the uptime falls below 90%, the system automatically routes your request to the best available alternative provider to ensure reliability. This protects your application from providers experiencing temporary issues. The fallback threshold (default `90%`) is configurable.
If the requested provider has low uptime but no alternative providers are
available for that model, the request will still be sent to the originally
requested provider.
#### Disabling Fallback with X-No-Fallback Header [#disabling-fallback-with-x-no-fallback-header]
If you need to bypass this protection and always use the exact provider you specified regardless of its current uptime, you can use the `X-No-Fallback` header:
```bash
# Force use of a specific provider even if it has low uptime
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-H "X-No-Fallback: true" \
-d '{
"model": "openai/gpt-4o",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
Using `X-No-Fallback: true` disables automatic provider failover. Your
requests will be sent to the specified provider even if it is experiencing
issues, which may result in higher error rates. Retries may still occur
against another key for the same provider when multiple keys are configured.
When the `X-No-Fallback` header is used, the routing metadata in logs will include `noFallback: true` to indicate that fallback was disabled for that request.
## Automatic Retry & Fallback [#automatic-retry--fallback]
When using model ID routing (without a provider prefix), Vichar automatically retries failed requests on alternate providers. This happens transparently within the same API call — your application receives the successful response as if nothing went wrong.
### How Retry Works [#how-retry-works]
1. Your request is routed to the best available provider using the smart routing algorithm
2. If that provider fails with a retryable error (see [What Triggers a Retry](#what-triggers-a-retry) below), the gateway marks the provider as failed
3. The next best available provider is selected and the request is retried
4. Up to **2 retries** by default (configurable per project via the routing configuration) are attempted before returning an error to the client
```
Request → Provider A (500 error) → Provider B (200 OK) → Response
```
Both streaming and non-streaming requests support automatic retry.
### What Triggers a Retry [#what-triggers-a-retry]
Retries are triggered by failures classified as **provider-side or gateway-side problems**:
* **5xx errors** (500 Internal Server Error, 502 Bad Gateway, 503 Service Unavailable, etc.)
* **Timeouts** (upstream provider took too long to respond)
* **Connection failures** (network errors, DNS failures, etc.)
* **Upstream rate limits** (a `429` from the provider)
* **Provider account and mapping problems** — an upstream `401`/`403` (bad provider credentials), `402` (provider account out of funds), `404`/`405` (model or endpoint mapping gap), and a few specific `400` bodies that indicate the same kinds of gateway-side problems
Retries are **not** triggered by:
* **4xx client errors** — a request that is genuinely invalid (validation errors, unsupported parameters) fails the same way everywhere, so it is passed through with its original status instead of retried
* **Content filter responses** (Azure ResponsibleAI, etc.)
### When Retry Is Disabled [#when-retry-is-disabled]
Automatic retry to a different provider is disabled when:
* The `X-No-Fallback: true` header is set
* A specific provider is requested (e.g., `openai/gpt-4o`)
* The request carries a session id and [sticky session routing](#sticky-session-routing) is enabled — the session stays pinned to its provider
* No alternative providers are available for the requested model
* The maximum retry count (default 2) has been exhausted
Retries can still happen within the same provider when multiple keys are
configured and the current key fails with a retryable error.
### Routing Transparency [#routing-transparency]
Every provider attempt — both failed and successful — is recorded in the `routing` array in the response metadata (streaming and non-streaming alike) and activity logs:
```json
{
"metadata": {
"routing": [
{
"provider": "openai",
"model": "gpt-4o",
"status_code": 500,
"error_type": "server_error",
"succeeded": false,
"credentialSource": "byok",
"apiKeyHash": "f029ee9",
"providerKeyId": "pk_2f9a...",
"providerKeyLabel": "billing-team-key"
},
{
"provider": "azure",
"model": "gpt-4o",
"status_code": 200,
"error_type": "none",
"succeeded": true,
"credentialSource": "platform",
"apiKeyHash": "ecb88d5"
}
]
}
}
```
#### Whose key served each attempt [#whose-key-served-each-attempt]
`credentialSource` says who owns the provider credential an attempt was sent with:
| Value | Meaning |
| ---------- | --------------------------------------------------------------------------------------------------------- |
| `byok` | Your own provider key. The provider bills you directly and the attempt is not deducted from your credits. |
| `platform` | An Vichar credential. The attempt runs on credits and is deducted from your balance. |
This matters most in **hybrid** mode, where a request that fails on your own key falls back to Vichar's credential: both attempts appear in the same `routing` array, and only `credentialSource` tells them apart — `apiKeyHash` is an opaque fingerprint that says two attempts used different keys, not which key was yours. The same value is stored on the log as `routingMetadata.usedCredentialSource` for the credential that ultimately served the request, and is shown as a **your key** / **Vichar key** badge in the dashboard's routing view.
#### Which of your keys ran [#which-of-your-keys-ran]
A `byok` attempt also carries the key itself: `providerKeyId`, and `providerKeyLabel` — the key as it is named on your [provider keys](https://app.vichar.io/dashboard) page (its name, or its masked token when it has none). So when several of your keys are configured for a provider and the gateway rotates between them, each attempt says which one it used instead of leaving you to decode a fingerprint.
Chat requests additionally record `routingMetadata.eligibleProviderKeys` on the log: your keys that were candidates for the provider that served the request, in selection order. It is omitted for credits-mode projects, which route on Vichar credentials, and for custom providers, whose keys are scoped by their own catalogue.
These fields describe **your** keys only. Vichar's own credentials — the ones
that serve credits-mode traffic — are never named: a `platform` attempt still
reports `credentialSource` and `apiKeyHash`, but never `providerKeyId` or
`providerKeyLabel`.
### Retried Log Tracking [#retried-log-tracking]
Each provider attempt creates its own log entry. Failed attempts that were retried are marked with:
* **`retried: true`** — indicates this failed request was retried on another provider
* **`retriedByLogId`** — the ID of the final successful log entry
This allows you to distinguish between unrecovered failures and failures that were transparently recovered via retry. In the dashboard, retried logs display a "Retried" badge with a link to the successful log.
### Impact on Provider Health [#impact-on-provider-health]
Failed attempts still count against the provider's uptime score, even when the request was successfully retried on another provider. This means:
* A provider that keeps failing will see its uptime score drop
* Only gateway and upstream errors count: requests rejected as client errors (invalid request bodies, unsupported parameters) are excluded from both the error count and the request total, so your own bad requests never mark a provider as down
* The exponential uptime penalty kicks in below 95% (see [Smart Routing Algorithm](#smart-routing-algorithm))
* Future requests are automatically routed away from unreliable providers
* Your application stays reliable without any code changes on your side
Automatic retry and fallback works together with smart routing to provide
self-healing behavior. Failing providers are automatically avoided, and your
requests are transparently recovered on reliable alternatives.
## Per-Project Routing Configuration (Enterprise) [#per-project-routing-configuration-enterprise]
All plans use observed token usage for cache pricing when sufficient history exists. On the **Enterprise plan**, you can override the settings listed below **per project** from the dashboard under **Project Settings → Routing**, including explicit cache-pricing assumptions. The 24-hour usage window and minimum sample requirements are fixed; the **History** settings control uptime, throughput, and latency metrics.
Overrides are merged on top of the defaults, so you only set the values you want to change. When a custom configuration is disabled, the project falls back to the defaults.
The following groups can be customized per project:
| Group | What it controls | Defaults |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Weights** | Relative importance of each scoring factor | `price 0.6`, `imagePrice 1.0`, `uptime 0.5`, `throughput 0.05`, `latency 0.025`, `cache 0` |
| **Thresholds** | Cache prompt size and pricing overrides, uptime-penalty threshold, exploration rate, and fallback metrics | `cachePromptTokens 5000`, `cacheHitRate 0.1`, `cacheOutputRatio 0.2` (coding: `0.9` / `0.02`; Chat: `0.5` / `0.1`), `uptimePenalty 95`, `defaultUptime 100`, `defaultLatency 1000`, `defaultThroughput 50`, `explorationRate 0.01` |
| **Retry** | Max cross-provider fallback attempts and the low-uptime reroute threshold | `maxRetries 2`, `lowUptimeFallbackThreshold 90` |
| **Timeouts** | Per-request time limits (end-to-end, streaming, non-streaming) — see [Request Timeouts](https://docs.vichar.io/features/timeouts). Capped at the infrastructure defaults — an override can only lower them | `gatewayMs 1,500,000`, `streamingMs 1,200,000`, `plainMs 600,000` |
| **History** | The metrics window and the time-decay tier boundaries and weights | `windowMinutes 60` (max 120), `tier1Minutes 1`, `tier2Minutes 5`, `tier1Weight 10`, `tier2Weight 3`, `tier3Weight 1` |
| **Sticky** | Stable-provider preference: on/off, TTL, hard-switch uptime floor, soft-switch score margin | `enabled true`, `ttlSeconds 3600`, `uptimeThreshold 85`, `scoreMargin 0.15` |
| **Session** | [Sticky session routing](#sticky-session-routing): on/off, pin TTL, re-pin uptime floor | `enabled true`, `ttlSeconds 3600`, `uptimeThreshold 85` |
| **Provider priorities** | Per-provider priority multipliers; set a provider to `0` to disable it for that project | `1` for every provider |
Per-project routing configuration requires the Enterprise plan. If you'd like
to tune routing for your workloads, contact us at [contact@vichar.io](mailto:contact@vichar.io).
## Optimized Auto Routing [#optimized-auto-routing]
Auto routing automatically selects the best model for your specific use case without you having to specify a model at all.
### Default behaviour [#default-behaviour]
By default, auto routing picks the cheapest of a small built-in set of models that can serve the request: it filters out models whose context window, capabilities (vision, tools, reasoning, structured output) or provider availability do not fit, then selects the cheapest of what remains. Larger prompts skip the smallest model.
```bash
# Let Vichar choose the optimal model
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Your request here..."}]
}'
```
### Free Models Only [#free-models-only]
When using auto routing, you can restrict the selection to only free models (models with zero input and output pricing) by setting the `free_models_only` parameter to `true`:
```bash
# Auto route to free models only
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Hello!"}],
"free_models_only": true
}'
```
Adding even a small amount of credits to your account (e.g., $10) will
immediately upgrade your free model rate limits from 5 requests per 10 minutes
to 20 requests per minute (free-model use still requires a verified email).
The `free_models_only` parameter only works with auto routing (`"model":
"auto"`). If no free models are available that meet your request requirements,
the API will return an error.
### Reasoning models only [#reasoning-models-only]
Just specify the `reasoning_effort` value and only a model which supports reasoning will be chosen. This parameter is not specific to the auto model.
```bash
# Auto route only to reasoning models
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Hello!"}],
"reasoning_effort": "medium"
}'
```
### Exclude Reasoning Models [#exclude-reasoning-models]
When using auto routing, you can exclude reasoning models from selection by setting the `no_reasoning` parameter to `true`. This is useful when you want faster responses or need to avoid the additional cost and latency of reasoning models:
```bash
# Auto route excluding reasoning models
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Hello!"}],
"no_reasoning": true
}'
```
The `no_reasoning` parameter only works with auto routing (`"model": "auto"`).
If no non-reasoning models are available that meet your request requirements,
the API will return an error.
Auto routing analyzes your payload and automatically chooses between
cost-effective models for simple requests and more powerful models for complex
or large-context requests.
### How It Works [#how-it-works]
1. **Request Analysis**: The system analyzes your request including message content, context size, and any special parameters
2. **Model Selection**: Based on the analysis, it selects the most appropriate model considering cost, performance, and capabilities
3. **Transparent Routing**: Your request is seamlessly routed to the chosen model and provider
4. **Optimized Response**: You receive the best possible response while maintaining cost efficiency
Auto routing decisions are transparent in your usage logs, so you can always
see which model was selected for each request.
## Smart Routing [#smart-routing]
`"model": "auto"` above is fixed: it always picks from the same built-in set, and its behaviour does not change. **Smart routing** is a separate model string, `"model": "smart"`, where you choose the candidate models and how they are ranked.
Under **Organization settings → Smart Routing** you choose up to 30 models from the [models catalogue](https://app.vichar.io/dashboard) that `"model": "smart"` may resolve to, plus the classifier that ranks them. Individual projects can override the organization default on their **Settings → Routing** page; a project without an override inherits it. Owners and organization admins can edit the organization default, project admins the project override.
Smart routing is available to every organization, including pay-as-you-go, while it is in beta. `"model": "smart"` fails rather than falling back: an organization that cannot use it gets a 403, and one that has not configured it gets a 400 naming the setting — it never degrades quietly into the `auto` candidate set.
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "smart",
"messages": [{"role": "user", "content": "Your request here..."}]
}'
```
The configured list is exhaustive. If no model on it can serve a request — for example an image request against a text-only list — the gateway returns a 400 rather than falling back to a model you did not allow. `free_models_only` narrows the list to its free models rather than replacing it, so a request parameter cannot route outside what the organization allowed.
#### Classifiers [#classifiers]
* **None** — pick the cheapest model on the list that can serve the request. This is the default behaviour, restricted to your models.
* **Jev (TypeSafe)** — rate the request before routing it.
With the Jev classifier, the gateway sorts your models by blended average price and splits them into Low, Medium and High price bands. The dashboard previews that split from catalogue list prices; the gateway ranks only the providers your project can actually use, at your own rates, so the live split can differ. Each request is then classified for task type, output type and difficulty, and served from the matching band; the classifier's own model preference breaks ties inside the band when it is confident enough. A request rated High also gets a larger default reasoning effort when the selected model supports reasoning and the request left room for it — thinking is drawn from the same `max_tokens` allowance as the answer, so the cheaper default stands on a tight budget rather than risking a response that spends its whole allowance thinking.
The classifier adds one short round trip before the upstream call. It fails open: if it times out or errors, the request is served by the cheapest eligible model on your list and the log records the fallback.
#### What it costs [#what-it-costs]
Smart routing itself carries no platform fee — you pay for the models it selects. A Jev classification is billed at the catalogue rate for [`jev-1.13.0`](https://app.vichar.io/dashboardsafe) on TypeSafe, which is priced on input tokens only, and works out to roughly $0.0001 per call. Each call is recorded as its own log entry against the same organization, project and API key as the request that triggered it, so the amount is visible in your activity feed and usage analytics rather than estimated. The request's own log entry also carries the charge in its `smartRouting.classifierCost` routing metadata.
The call runs on our credential, so it is billed as credits even for a project using its own provider keys. Nothing is charged for the `None` classifier, for a verdict reused from a sticky session, or for a classifier call that fails.
**Sticky sessions classify once.** When a request carries a session id (see [sticky session routing](#sticky-session-routing)) and the project has it enabled, the first turn is classified and the rest of the session reuses that verdict and the model it resolved to — so a conversation is not re-rated on every turn and does not migrate between models mid-thread, which would cost it the upstream prompt cache. The pin expires with the session TTL. If the pinned model stops being available, the stored verdict is re-applied to the remaining candidates rather than triggering a fresh classification.
#### Routing metadata [#routing-metadata]
Every smart-routed request made against a configured list records a `smartRouting` block on its log entry, visible in the request detail view: the classifier used, the eligible and surviving candidate models, the difficulty, task and output type, the classifier's preferred model and confidence, the band that was served, the selected model, the classifier latency and cost, whether the classifier failed open, and whether the verdict was reused from earlier in the session.
## Best Practices [#best-practices]
### For Development [#for-development]
* Use specific model names during development and testing
* Leverage auto routing for production workloads to optimize costs
### For Production [#for-production]
* Use auto routing (`"model": "auto"`) for the best balance of cost and performance, or smart routing (`"model": "smart"`) to route across models you choose
* Monitor your usage patterns through the dashboard to understand routing decisions
* Set up provider keys for multiple providers to maximize routing options
### For Cost Optimization [#for-cost-optimization]
* Let auto routing handle model selection to automatically use the most cost-effective options, or smart routing to spend more only on the requests that need it
* Use model IDs without provider prefixes to always get the cheapest available provider
* Monitor your usage analytics to track cost savings from intelligent routing
# Service Tiers
URL: https://docs.vichar.io/features/service-tiers
Some OpenAI, Google, and Fireworks models support selectable **processing tiers** that trade
latency and availability against price. You pick one per request with the
OpenAI-compatible `service_tier` parameter, and Vichar forwards it only
when the selected provider/model mapping supports that tier.
| Tier | `service_tier` | Cost vs. standard | Latency / availability |
| ------------ | ------------------------- | ----------------- | ------------------------------------------- |
| Standard | `default` / `auto` / omit | baseline | Normal on-demand latency |
| **Flex** | `flex` | **−50%** | Best-effort; may be preempted under load |
| **Priority** | `priority` | varies by model | Prioritized above standard and flex traffic |
## Using the `service_tier` parameter [#using-the-service_tier-parameter]
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google-vertex/gemini-3.1-pro-preview",
"service_tier": "priority",
"messages": [
{ "role": "user", "content": "Summarize this incident report." }
]
}'
```
Accepted values are `flex`, `priority`, and `default`/`auto` (standard). If you
request `flex` or `priority` for a provider/model mapping that does not support
that tier, the gateway returns a 400 `unsupported_service_tier` error and logs
the request as a client error.
The parameter works the same on the OpenAI-compatible **Responses API**
(`/v1/responses`): the tier is forwarded to the provider and the response's
`service_tier` field echoes the tier that was actually served. The **Images
API** (`/v1/images/generations` and `/v1/images/edits`) accepts it too and
forwards it to the underlying image model, so image mappings that list Flex
can be generated at the Flex rate.
Coding (dev) plans are limited to `default`/`auto` and `flex` — the premium
`priority` tier is not part of the plan and a request that asks for it returns a 403.
## Supported providers [#supported-providers]
Service tiers are explicit per provider/model mapping. Check the model page for
the exact tiers exposed by each provider card.
* **OpenAI** (`openai`) — sent as the OpenAI `service_tier` request field for
supported OpenAI models. Flex is billed at 0.5x standard token prices and
Priority uses the model-specific multiplier shown on the model page.
* **Google Vertex AI** (`google-vertex`) — sent as the
`X-Vertex-AI-LLM-Shared-Request-Type` request header, together with
`X-Vertex-AI-LLM-Request-Type: shared` so the request bypasses any Provisioned
Throughput on the project and actually reaches the shared Flex/Priority tier.
Flex and Priority are served only on the **global** endpoint, which is the
gateway default. Google Flex PayGo applies a 0.5x multiplier; Google Priority
PayGo applies a 1.8x multiplier.
* **Google AI Studio / Gemini API** (`google-ai-studio`) — sent as a
`service_tier` field in the request body for configured models that opt in.
* **Fireworks AI** (`fireworks`) — sent as a `service_tier` field in the request
body on its OpenAI-compatible chat completions endpoint. Only Priority is
offered (Fireworks publishes no Flex rate card), billed at a 1.25x multiplier.
Fireworks does not report the tier it served, but a Priority request is either
served at Priority or shed with a 503, so an accepted request is billed at the
tier it was sent at.
* **Azure** (`azure`) — sent as a `service_tier` field in the request body on the
v1 chat completions and responses endpoints. Only Priority is offered (Azure
publishes no Flex rate card); the multiplier is model-specific and shown on the
model page. Azure requires a **Global Standard** or **Data Zone (US)**
deployment of a model version `2025-12-01` or later, and a subscription
entitled to priority processing. Which deployment types a given model offers
the tier on differs per model — some are Global Standard only — so check
Microsoft's [priority processing
availability](https://learn.microsoft.com/azure/foundry/openai/concepts/priority-processing)
tables for the model you deploy. Azure downgrades to standard when the
entitlement is missing, during peak load, on ramp-rate limits, and for
long-context prompts on some models; it reports the served tier back, so a
downgraded request is billed at standard and reported as such.
Tiers are supported on a **subset** of models, and the Flex and Priority
subsets differ by provider. For example, Google Flex PayGo lists Gemini 3
image / Nano Banana models, but Google Priority PayGo does not; those
configured image mappings are Flex-only.
Flex and Priority are only honored when the request reaches the provider
directly, so a provider key with a **custom base URL** (a proxy) is excluded
from service-tier routing — a proxy may silently drop the tier and serve
standard. This applies to every provider that offers tiers, including OpenAI:
the tier travels as a `service_tier` body field that an OpenAI-compatible
proxy is free to ignore. With multiple providers/keys, the gateway routes
around the ineligible key automatically; if a request pins a provider whose
only key uses a custom base URL, it returns a 400 instead of silently
downgrading. Keys with no custom base URL (the managed default) are always
eligible.
## Retries and fallback never downgrade the tier [#retries-and-fallback-never-downgrade-the-tier]
A request can change provider or credential mid-flight: Vichar falls back
to another provider when one fails, and rotates to another key for the same
provider when a credential returns a 429 or an auth error. A requested tier is
carried through all of it.
* Provider routing is narrowed to mappings that support the requested tier
**before** a provider is picked, so no fallback candidate can be one that
would serve the request as standard.
* Key selection — BYOK keys, platform-managed credentials, and env credentials
alike — skips any credential that cannot carry the tier (a proxy base URL, or
a Vertex credential pinned to a regional endpoint), on the first attempt and
on every retry.
* Every attempt re-resolves the tier against the provider, region and
credential it actually resolved to, and fails rather than sending at a lower
tier. If no eligible candidate is left, you get the upstream error instead of
a silently downgraded response.
This applies to a tier you requested yourself. The optional coding-plan default
tier is a cost preference rather than a requirement, so a request that cannot be
served at that tier runs at standard instead of failing; `used_service_tier` in
the response metadata always reports what was actually served. A tier sent on
the request always takes precedence over that default, within the tiers the plan
allows.
Rotating to another key for the same provider is a separate upstream account
as far as prompt caching is concerned, so a retried request re-writes its
cached prefix rather than reading the original one. Send `x-no-fallback: true`
to have the original upstream error returned to you — note that this disables
cross-provider fallback, not key rotation within a provider.
## Pricing uses multipliers [#pricing-uses-multipliers]
Service tiers do not define separate model prices in Vichar. They multiply
the provider mapping's standard token prices:
* Standard / `default` / `auto`: 1x
* Flex: 0.5x
* Priority: model/provider-specific, shown on the model page
The multiplier scales per-token costs, including input, output, cached, and
image tokens. Flat per-request and web-search fees are not tier-scaled.
## Billing follows the served tier [#billing-follows-the-served-tier]
When a provider reports the tier that was actually served, Vichar bills
that returned tier instead of blindly billing the requested value:
* A `priority` request that runs as priority is billed at that provider mapping's priority multiplier, shown on the [model page](https://app.vichar.io/dashboard).
* A `flex` request that runs as flex is billed at 0.5x.
* A request that is served as standard is billed at the standard 1x rate.
The served tier is read back from the provider response — Vertex reports it in
`usageMetadata.trafficType` (`ON_DEMAND_PRIORITY` / `ON_DEMAND_FLEX` /
`ON_DEMAND`), Google AI Studio reports it in the `x-gemini-service-tier`
response header, and OpenAI can return `service_tier` in response payloads or
stream events. Providers that report no tier at all (Fireworks) never downgrade
silently — they reject the request instead — so an accepted request is billed at
the tier it was sent at.
Vichar rejects unsupported tier requests before provider routing. For example,
`gemini-3-pro-image` exposes Flex and Priority for Google AI Studio, but only
Flex for Vertex. Other mappings expose neither tier.
You can see per-tier pricing for each model on its
[model page](https://app.vichar.io/dashboard). Supported provider cards include a
Service Tier selector in the card header and show the active multiplier next to
each tier.
## Sources [#sources]
* [OpenAI API pricing](https://openai.com/api/pricing/)
* [Google Flex PayGo](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/flex-paygo)
* [Google Priority PayGo](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/priority-paygo)
* [Fireworks Serverless Priority and Fast](https://docs.fireworks.ai/serverless/priority-and-fast)
# Source Attribution
URL: https://docs.vichar.io/features/source
The `X-Source` header allows you to identify your domain when making requests to Vichar. This information is used to generate public usage statistics showing how Vichar is being used across different websites and applications.
## X-Source Header [#x-source-header]
Include the `X-Source` header with your domain name in your requests:
```bash
curl -X POST https://api.vichar.io/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "X-Source: example.com" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "Hello, how are you?"
}
]
}'
```
## Domain Format [#domain-format]
The `X-Source` header accepts domain names in various formats. All of the following are valid and will be normalized to the same domain:
* `example.com`
* `https://example.com`
* `https://www.example.com`
* `www.example.com`
All variations will be stripped down to the base domain (`example.com`) for aggregation purposes.
After stripping the protocol and `www.` prefix, the value must contain only alphanumeric characters, hyphens, dots, and slashes — a value with other characters (underscores, spaces, query strings) fails the whole request with a `400`.
When no `X-Source` header is sent, the gateway falls back to the `HTTP-Referer` header, and recognizes well-known coding agents from the `User-Agent` header for source attribution.
## Public Statistics [#public-statistics]
Data from the `X-Source` header is used to generate public statistics about Vichar usage, including:
* **Popular Domains**: Which websites and applications are using Vichar most frequently
* **Model Usage**: What models are being used by different domains
* **Geographic Distribution**: Where requests are coming from across different sources
* **Growth Trends**: How usage is growing over time for different domains
These statistics help demonstrate the adoption and impact of Vichar across the ecosystem.
## Privacy Considerations [#privacy-considerations]
### What's Public [#whats-public]
* Domain names (stripped of protocol and www prefixes)
* Aggregated request counts and model usage
* General geographic regions (country-level data)
### What's Private [#whats-private]
* Individual request content or responses
* User identifiers or personal information
* Detailed usage patterns beyond aggregated counts
* API keys or authentication details
## Benefits [#benefits]
Including the `X-Source` header provides several benefits:
### For Your Project [#for-your-project]
* **Recognition**: Your domain will appear in public usage statistics
* **Credibility**: Demonstrates real-world usage of your application
* **Community**: Contributes to the broader Vichar ecosystem
### For the Community [#for-the-community]
* **Transparency**: Shows real adoption and usage patterns
* **Inspiration**: Other developers can see successful implementations
* **Growth**: Helps demonstrate the value of open-source LLM infrastructure
## Optional but Recommended [#optional-but-recommended]
While the `X-Source` header is optional, we strongly encourage its use to:
* Support transparency in the Vichar ecosystem
* Help showcase successful integrations
* Contribute to understanding of LLM usage patterns
* Demonstrate the real-world impact of your application
Your participation helps build a more transparent and collaborative LLM ecosystem.
# Speech Generation
URL: https://docs.vichar.io/features/speech-generation
Vichar supports text-to-speech (TTS) through the OpenAI-compatible
**`/v1/audio/speech`** endpoint, powered by ElevenLabs, Google Gemini, OpenAI,
and Alibaba Qwen speech models.
## Available Models [#available-models]
Browse all speech generation models, with up-to-date pricing, on the
[models page](https://app.vichar.io/dashboard).
Billing varies by model family. Some models are billed on token usage reported
by the provider (input text tokens and output audio tokens), while others are
billed on input character count (those return audio bytes without usage data).
See the [models page](https://app.vichar.io/dashboard) for each model's exact
pricing.
## Parameters [#parameters]
| Parameter | Type | Default | Description |
| ----------------- | ------ | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `model` | string | required | The speech model to use |
| `input` | string | required | The text to synthesize into speech |
| `voice` | string | model | A prebuilt voice. Defaults to `Kore` (Gemini), `alloy` (OpenAI), `Sarah` (ElevenLabs), or the model's first voice on Qwen (`longanlingxin` on Plus, `longanhuan_v3.6` on Flash) |
| `response_format` | string | model | Audio format. OpenAI: `mp3` (default), `opus`, `aac`, `flac`, `wav`, `pcm`. ElevenLabs: `mp3` (default), `wav`, `pcm`, `opus`. Gemini: `wav` (default), `pcm`. Qwen: `wav` |
| `instructions` | string | — | Optional style/delivery directive prepended to the input (e.g. `"Say cheerfully"`) |
| `speed` | number | — | Accepted for OpenAI compatibility, but not applied by Gemini speech models |
Gemini speech models return raw PCM audio. Vichar wraps it in a WAV container
by default (`response_format: "wav"`), or returns the raw 16-bit little-endian
PCM at 24 kHz when `response_format: "pcm"` is requested. Other formats
such as `mp3` are only available on the OpenAI models, which return the audio
already encoded in the requested format.
## curl [#curl]
```bash
curl -X POST "https://api.vichar.io/v1/audio/speech" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-2.5-flash-preview-tts",
"input": "Hello, welcome to Vichar!",
"voice": "Kore"
}' \
--output speech.wav
```
## OpenAI SDK [#openai-sdk]
Works with the standard OpenAI client library — just point the base URL to
Vichar.
```ts
import OpenAI from "openai";
import { writeFileSync } from "fs";
const openai = new OpenAI({
apiKey: process.env.LLM_GATEWAY_API_KEY,
baseURL: "https://api.vichar.io/v1",
});
const response = await openai.audio.speech.create({
model: "gemini-2.5-flash-preview-tts",
voice: "Kore",
input: "Hello, welcome to Vichar!",
});
const buffer = Buffer.from(await response.arrayBuffer());
writeFileSync("speech.wav", buffer);
```
## Streaming [#streaming]
Streaming speech responses (chunked audio or `stream_format: "sse"`) are not
supported yet. The endpoint always returns the complete audio file in a single
response, so there is no low-latency, play-as-you-go output for now.
## Voices [#voices]
Gemini exposes 30 prebuilt voices. A few common ones:
`Kore`, `Puck`, `Zephyr`, `Charon`, `Fenrir`, `Leda`, `Orus`, `Aoede`. When
`voice` is omitted on a Gemini model, `Kore` is used.
OpenAI voices include `alloy`, `ash`, `ballad`, `coral`, `echo`, `fable`,
`nova`, `onyx`, `sage`, `shimmer`, and `verse`. When `voice` is omitted on an
OpenAI model, `alloy` is used.
ElevenLabs models accept 20 named voices, including `Sarah`, `Aria`, `Roger`,
`Laura`, `Charlie`, `George`, `Charlotte`, `Jessica`, `Brian`, and `Lily`. When
`voice` is omitted on an ElevenLabs model, `Sarah` is used. A raw ElevenLabs
voice id is also accepted directly.
Qwen-Audio-3.0-TTS voices are model-specific and cannot be mixed between
models: `longanlingxin` and `longanlufeng` on Plus (default `longanlingxin`),
and `longanhuan_v3.6`, `longjielidou_v3.6`, `loongeva_v3.6`, and `loongjohn` on
Flash (default `longanhuan_v3.6`).
## ElevenLabs [#elevenlabs]
The four ElevenLabs models are billed per **input character** (see the [models
page](https://app.vichar.io/dashboard) for rates):
* `eleven-multilingual-v2` — most lifelike, rich emotional expression, 29 languages
* `eleven-v3` — most expressive and human-like, 70+ languages
* `eleven-flash-v2-5` — ultra-low latency, 32 languages
* `eleven-turbo-v2-5` — fast and balanced, 32 languages
```bash
curl -X POST "https://api.vichar.io/v1/audio/speech" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "eleven-multilingual-v2",
"input": "Hello, welcome to Vichar!",
"voice": "Sarah"
}' \
--output speech.mp3
```
# System One
URL: https://docs.vichar.io/features/system-one
Vichar exposes a `/v1/systemone` endpoint for **System One** models: models
that read a state, answer questions you name, and return typed decisions with
calibrated probabilities. There is no free-form text in the response — every
answer is a value your code can branch on directly.
Use it when a model call exists only to make a decision: routing a support
ticket, scoring a retrieved passage, checking whether a citation supports a
claim, or classifying a record. A chat model can do these too, but you then have
to parse prose, and you get no probability to threshold on.
Browse available decision models on the
[models page](https://app.vichar.io/dashboard).
For the full request and response schema, see the
[API reference](https://docs.vichar.io/v1_systemone).
## Endpoint [#endpoint]
`POST https://api.vichar.io/v1/systemone`
## Question types [#question-types]
Every question has a `type`, `instructions`, and — for choice and score — its
own `criteria`. Answers come back under the ids you chose.
| Type | Ask | Answer |
| -------- | -------------------------------- | -------------------------------------------------------------------------- |
| `noul` | A yes/no question | `noul`: probability the answer is yes, 0 to 1 |
| `choice` | One option from a set you define | `choice`, `probabilities` per option, `confidence` |
| `score` | A rating across ordered levels | `score` (can land between levels), `probabilities`, `legend`, `confidence` |
## cURL [#curl]
```bash
curl -X POST "https://api.vichar.io/v1/systemone" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-1.13.0",
"state": "Help! My payouts have been failing for 3 days.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations",
"sales": "Pricing, upgrades, new accounts"
}
},
"is_urgent": {
"type": "noul",
"instructions": "Does this convey urgency?"
},
"impact": {
"type": "score",
"instructions": "Rate the operational impact.",
"criteria": ["None", "Limited", "Critical"]
}
}
}'
```
```json
{
"model": "typesafe/jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 0.88, "technical": 0.12, "sales": 0.0 },
"confidence": 0.81
},
"is_urgent": { "type": "noul", "noul": 0.95 },
"impact": {
"type": "score",
"score": 1.9,
"legend": { "0": "None", "1": "Limited", "2": "Critical" },
"probabilities": { "0": 0.0, "1": 0.1, "2": 0.9 },
"confidence": 0.88
}
},
"usage": { "input_tokens": 318, "output_tokens": 34 }
}
```
The `model` field reports the pinned model that answered, prefixed with the
provider it was served by. Provider aliases that move between releases (for
example `jev-latest`) are accepted and resolve to the pinned version, so a
request is always billed and logged against the version you can read prices for.
## Request fields [#request-fields]
| Field | Type | Description |
| ----------- | --------------------------- | -------------------------------------------------------------------------- |
| `model` | `string` | Decision model to use. Optionally prefixed with a provider (`typesafe/…`). |
| `state` | `string \| object \| array` | The content to evaluate. Structured data lets questions reference fields. |
| `questions` | `map` | At least one typed question, keyed by ids you choose. |
## Deciding in code [#deciding-in-code]
The point of a typed answer is that the policy stays in your code, not in a
prompt. Threshold the probability, and use `confidence` as a second axis to send
uncertain cases to a human instead of acting on a coin flip:
```ts
const res = await fetch("https://api.vichar.io/v1/systemone", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "jev-1.13.0",
state: ticket,
questions: {
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: { billing: null, technical: null, sales: null },
},
is_urgent: { type: "noul", instructions: "Does this convey urgency?" },
},
}),
});
const { answers } = await res.json();
if (answers.department.confidence < 0.7) {
await queueForTriage(ticket);
} else {
await route(ticket, answers.department.choice, {
priority: answers.is_urgent.noul > 0.8 ? "high" : "normal",
});
}
```
Ask everything you might need in one request: the state is read once and every
question is evaluated against it, so a batch of questions costs far less than
one request each — including speculative questions you only read when another
answer makes them relevant.
Decision models are billed on input tokens only — the state plus every
question. Output tokens are free, since the response is a set of values rather
than generated text. Question text counts as input on every call, so a large
rubric has a real per-request cost.
Decision models only work on `/v1/systemone`. Requesting one on
`/v1/chat/completions` returns a 400 pointing you at the right endpoint, they
cannot generate text or call tools, and they are not available in the
playground. Input is text only.
# Request Timeouts
URL: https://docs.vichar.io/features/timeouts
The gateway enforces hard time limits on every request. These protect the
platform from stuck upstream connections, but they matter to you directly if
you run long generations or agentic pipelines: a single completion that runs
longer than the limit is terminated.
| Limit | Default | Applies to |
| ----------------- | ---------- | -------------------------------------------------------- |
| **Streaming** | 20 minutes | Streaming completions (`stream: true`) |
| **Non-streaming** | 10 minutes | Non-streaming completions |
| **End-to-end** | 25 minutes | Overall ceiling for the whole request, including retries |
## How the limits behave [#how-the-limits-behave]
* **They are total-duration limits, not idle limits.** The timer starts when
the gateway opens the upstream provider request and keeps running for the
entire response — a stream is cut at the limit even while it is actively
producing tokens.
* **They apply per request, not per session or conversation.** Every
completion call gets a fresh window. A long-running agent that makes many
calls over hours is unaffected; only a *single* call exceeding the limit
fails.
* **Tool execution doesn't count.** When a completion finishes with
`tool_calls`, the HTTP request ends. The time your application spends
running tools (or coordinating subagents) between requests never consumes
the window.
* **On timeout** the gateway returns a `504` with type `timeout_error` (see
[Error Handling](https://docs.vichar.io/resources/error-handling)). If the limit is hit
mid-stream, the stream terminates.
## Long-running agentic workloads [#long-running-agentic-workloads]
Coordinator/subagent architectures often hold a "master" completion open while
work happens elsewhere. To stay within the limits:
* **Keep any single completion under 20 minutes.** Structure the coordinator
as a loop of discrete completions (the standard tool-calling pattern) rather
than one long-lived request that spans the whole investigation.
* **Stream long generations.** Non-streaming requests are capped at 10
minutes; streaming raises the per-request budget to 20 minutes and delivers
partial output as it is produced.
* **Split very long outputs.** If a single generation legitimately needs more
than 20 minutes (very large outputs on slow models, extensive reasoning),
break it into continuation requests.
## Changing the limits [#changing-the-limits]
**Per-project overrides (Enterprise)** can be set under **Project Settings →
Routing** — see [Per-Project Routing
Configuration](https://docs.vichar.io/features/routing#per-project-routing-configuration-enterprise).
Overrides can only *lower* the timeouts: the defaults above are the
infrastructure ceiling on the hosted platform and cannot be raised per
project.
**Self-hosted deployments** control the limits with environment variables on
the gateway service:
| Variable | Default | Controls |
| ------------------------- | --------- | ------------------------- |
| `AI_STREAMING_TIMEOUT_MS` | `1200000` | Streaming completions |
| `AI_TIMEOUT_MS` | `600000` | Non-streaming completions |
| `GATEWAY_TIMEOUT_MS` | `1500000` | End-to-end ceiling |
When raising the limits on a self-hosted deployment, raise the surrounding
infrastructure in lockstep: your load balancer's backend/response timeout and
the gateway's shutdown grace period (`SHUTDOWN_GRACE_PERIOD_MS`, and e.g.
Kubernetes `terminationGracePeriodSeconds`) must all be at least as long as
the longest stream you allow, or rollouts and intermediaries will still cut
long requests.
If your workload genuinely needs single completions longer than 20 minutes on
the hosted platform, contact us at [contact@vichar.io](mailto:contact@vichar.io) — the ceiling is an
infrastructure setting, not a per-model constraint.
# Transcription
URL: https://docs.vichar.io/features/transcription
Vichar exposes a dedicated `/v1/audio/transcriptions` endpoint for
speech-to-text. It transcribes audio files into text with word-level
timestamps, optional speaker diarization, and inverse text normalization
(spoken numbers and currencies formatted in their written form).
Use it when you want to:
* Turn recordings, voicemails, or podcast episodes into text
* Generate captions or searchable transcripts with per-word timing
* Feed spoken content into a downstream model or RAG pipeline
For the full request and response schema, see the
[API reference](https://docs.vichar.io/v1_audio_transcriptions).
## Endpoint [#endpoint]
`POST https://api.vichar.io/v1/audio/transcriptions`
Authenticate with your Vichar API key:
```bash
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY"
```
The current model is `grok-stt-1-0`, billed at **$0.10 per hour** of input
audio (against the duration reported by the provider).
## Parameters [#parameters]
The request body is `multipart/form-data`. Either `file` or `url` must be
provided.
| Parameter | Type | Default | Description |
| -------------- | ------ | -------- | --------------------------------------------------------------------------------------------------------- |
| `model` | string | required | The transcription model to use |
| `file` | file | — | The audio file to transcribe (WAV, MP3, OGG, Opus, FLAC, AAC, MP4, M4A, and more) |
| `url` | string | — | URL of an audio file to download and transcribe instead of uploading one |
| `language` | string | — | Language code (e.g. `en`). When set, enables formatting of numbers and currencies into their written form |
| `diarize` | string | `false` | When `"true"`, each word in the response includes a `speaker` field identifying the detected speaker |
| `filler_words` | string | `false` | When `"true"`, filler words (e.g. "uh", "um") are kept in the transcript instead of being removed |
| `keyterm` | string | — | A key term to bias transcription toward (e.g. product names). Repeat the field for multiple terms |
## Response [#response]
The response includes the full transcript, audio duration, and word-level
timestamps:
```json
{
"text": "The balance is $167,983.15.",
"language": "English",
"duration": 3.45,
"words": [
{ "text": "The", "start": 0.24, "end": 0.48 },
{ "text": "balance", "start": 0.48, "end": 0.96 },
{ "text": "is", "start": 0.96, "end": 1.12 },
{ "text": "$167,983.15.", "start": 1.12, "end": 3.2 }
]
}
```
## curl [#curl]
```bash
curl -X POST "https://api.vichar.io/v1/audio/transcriptions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-F model=grok-stt-1-0 \
-F language=en \
-F file=@audio.mp3
```
### Transcribing from a URL [#transcribing-from-a-url]
```bash
curl -X POST "https://api.vichar.io/v1/audio/transcriptions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-F model=grok-stt-1-0 \
-F url="https://example.com/audio.mp3"
```
# Video Generation
URL: https://docs.vichar.io/features/video-generation
Vichar supports asynchronous video generation through an OpenAI-compatible `POST /v1/videos` flow.
Currently available models:
* **Veo 3.1** through `google-vertex` (720p, 1080p, 4k)
* **Seedance 2.5** through `bytedance` (480p, 720p, 1080p; clips up to 30 seconds)
* **Seedance 2.0** and **Seedance 1.5 Pro** through `bytedance` (720p, 1080p), **Seedance 2.0 Fast** through `bytedance` (720p only)
* **Seedance 2.0 Mini** through `bytedance` (480p, 720p)
* **KLING v3.0** and **KLING v3.0 Turbo** through `atlascloud` (720p, 1080p; KLING v3.0 also 4k)
* **MiniMax H3 Max** through `minimax` (480p, 768p; clips from 5 to 15 seconds, always with audio)
You can find the current list of video-capable models on our [models page with the video filter enabled](https://app.vichar.io/dashboard) or programmatically through the [/v1/models endpoint](https://docs.vichar.io/v1_models).
## What Works Today [#what-works-today]
* `POST /v1/videos`
* `GET /v1/videos/{video_id}`
* `GET /v1/videos/{video_id}/content`
* Optional signed callbacks with `callback_url` and `callback_secret`
## Request Format [#request-format]
Vichar currently supports a focused subset of the OpenAI video API.
### Supported fields [#supported-fields]
| Field | Type | Required | Description |
| ------------------ | ------- | -------- | -------------------------------------------------------------------------------------------------------------------------- |
| `model` | string | yes | Any video-capable model from the filtered models page |
| `prompt` | string | yes | Text prompt for the video |
| `seconds` | number | yes | Duration in seconds. Supported values depend on the model (see below) |
| `size` | string | no | `widthxheight`, limited to the sizes supported by the selected model and provider |
| `audio` | boolean | no | Whether to include audio in the output (default `true`). Only honored when the model supports both audio and silent output |
| `image` | object | no | Optional first frame for image-to-video generation (see first/last frame inputs below) |
| `last_frame` | object | no | Optional ending frame when `image` is provided (see first/last frame inputs below) |
| `reference_images` | array | no | One to three provider-specific image inputs |
| `input_reference` | object | no | Alias for one or more `reference_images` |
| `reference_videos` | array | no | One to three reference video HTTPS URLs (Seedance 2.x only, see below) |
| `reference_audios` | array | no | One to three reference audio HTTPS URLs (Seedance 2.x only, see below) |
| `callback_url` | string | no | Vichar extension for completion webhooks |
| `callback_secret` | string | no | Vichar extension used to sign webhook deliveries |
### Sizes and durations by model [#sizes-and-durations-by-model]
| Model family | Provider | Supported sizes | Supported durations |
| ----------------- | --------------- | --------------------------------------------------------------------------------- | ------------------- |
| Veo 3.1 | `google-vertex` | `1280x720`, `720x1280`, `1920x1080`, `1080x1920`, `3840x2160`, `2160x3840` | `4`, `6`, `8`, `10` |
| Seedance 2.5 | `bytedance` | `848x480`, `854x480`, `480x854`, `1280x720`, `720x1280`, `1920x1080`, `1080x1920` | `4`–`30` |
| Seedance 2.0 | `bytedance` | `1280x720`, `720x1280`, `1920x1080`, `1080x1920` | `4`–`15` |
| Seedance 2.0 Fast | `bytedance` | `1280x720`, `720x1280` | `4`–`15` |
| Seedance 2.0 Mini | `bytedance` | `1280x720`, `720x1280`, `848x480`, `854x480`, `480x854` | `4`–`15` |
| Seedance 1.5 Pro | `bytedance` | `1280x720`, `720x1280`, `1920x1080`, `1080x1920` | `5`, `10` |
| KLING v3.0 | `atlascloud` | `1280x720`, `720x1280`, `1920x1080`, `1080x1920`, `3840x2160`, `2160x3840` | `5`, `10` |
| KLING v3.0 Turbo | `atlascloud` | `1280x720`, `720x1280`, `1920x1080`, `1080x1920` | `5`, `10` |
| MiniMax H3 Max | `minimax` | `848x480`, `854x480`, `480x854`, `1366x768`, `768x1366` | `5`–`15` |
Requests return `400` when the selected provider cannot serve the requested `size` or `seconds`. Seedance and KLING v3.0 derive `aspect_ratio` from the requested `size` (16:9 for landscape, 9:16 for portrait).
**KLING v3.0 Turbo** and **MiniMax H3 Max** always generate audio and do not support `audio: false`; silent requests return a `400`. Use **KLING v3.0** (`kling-v3-0`) when you need silent output.
### First/last frame inputs [#firstlast-frame-inputs]
Frame inputs interpolate a video between a starting frame and an optional ending frame. You provide the first frame as `image` and, optionally, the ending frame as `last_frame`. The gateway tags each one with the correct role for the provider, so you don't set roles yourself.
| Field | Required | Accepted input | Available on |
| ------------ | ------------------ | -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| `image` | for frame mode | HTTPS URL **or** base64 data URL | **Seedance 2.x** (`bytedance`), **KLING v3.0 / Turbo** (`atlascloud`), Veo 3.1 (`google-vertex`), `minimax`, `xai` |
| `last_frame` | no (needs `image`) | HTTPS URL **or** base64 data URL | **Seedance 2.x** (`bytedance`), **KLING v3.0 / Turbo** (`atlascloud`), Veo 3.1 (`google-vertex`), **MiniMax H3 Max** (`minimax`) |
#### Rules and limits [#rules-and-limits]
* **Seedance scope.** Frame inputs are supported on **Seedance 2.5**, **Seedance 2.0**, **Seedance 2.0 Fast**, and **Seedance 2.0 Mini** (`seedance-2-5`, `seedance-2-0`, `seedance-2-0-fast`, `seedance-2-0-mini`). Sending `image`/`last_frame` to Seedance 1.5 Pro or any other ByteDance model returns a `400`.
* **KLING scope.** Frame inputs are supported on **KLING v3.0** and **KLING v3.0 Turbo** (`kling-v3-0`, `kling-v3-0-turbo`). The gateway uploads base64 frames to AtlasCloud's media endpoint automatically, so both HTTPS URLs and base64 data URLs are accepted.
* **`last_frame` requires `image`.** Providing `last_frame` without `image` returns a `400`.
* **Not combinable with references.** First/last frame inputs (`image`, `last_frame`) cannot be combined with reference inputs (`reference_images`, `input_reference`, `reference_videos`, `reference_audios`).
#### Example (Seedance 2.0) [#example-seedance-20]
```bash
curl -X POST "https://api.vichar.io/v1/videos" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-0",
"prompt": "Morph smoothly from the first frame into the last frame",
"seconds": 5,
"size": "1280x720",
"image": { "image_url": "https://example.com/first-frame.png" },
"last_frame": { "image_url": "https://example.com/last-frame.png" }
}'
```
#### Example (KLING v3.0 image-to-video) [#example-kling-v30-image-to-video]
```bash
curl -X POST "https://api.vichar.io/v1/videos" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kling-v3-0",
"prompt": "Animate the scene with gentle camera motion",
"seconds": 5,
"size": "1280x720",
"image": { "image_url": "https://example.com/first-frame.png" }
}'
```
### Reference-guided generation (Seedance 2.x) [#reference-guided-generation-seedance-2x]
Seedance 2.x (`seedance-2-5`, `seedance-2-0`, `seedance-2-0-fast`, `seedance-2-0-mini`) can generate a video that is guided by reference **images**, **videos**, and **audio** — sometimes called omni-reference. You attach references as top-level fields in the same `POST /v1/videos` payload; the gateway forwards each one to the provider tagged with the correct role, so you don't set roles yourself.
| Reference type | Payload field | Count | Accepted input | Available on |
| -------------- | -------------------------------------------- | ----- | -------------------------------- | --------------------------------------- |
| Image | `reference_images` (`input_reference` alias) | 1–3 | HTTPS URL **or** base64 data URL | Seedance 2.x, Veo 3.1 (`google-vertex`) |
| Video | `reference_videos` | 1–3 | HTTPS URL only | Seedance 2.x |
| Audio | `reference_audios` | 1–3 | HTTPS URL only | Seedance 2.x |
Each list item accepts either a bare URL string or an object form:
* `reference_images`: `"https://…/subject.png"` or `{ "image_url": "https://…/subject.png" }`
* `reference_videos`: `"https://…/motion.mp4"` or `{ "video_url": "https://…/motion.mp4" }`
* `reference_audios`: `"https://…/track.mp3"` or `{ "audio_url": "https://…/track.mp3" }`
You can mix all three reference types in one request. The `prompt` can be a light instruction (for example `"adapt this to show more detail"`) — the references drive the result.
#### Rules and limits [#rules-and-limits-1]
* **HTTPS only for video and audio.** `reference_videos` and `reference_audios` must be publicly reachable HTTPS URLs (the provider fetches them). base64 data URLs are rejected for video/audio; images may be HTTPS URLs or base64 data URLs.
* **Reference video resolution.** Seedance requires reference video frames to be at least \~409,600 pixels (roughly 480p or larger). Low-resolution clips such as 360p are rejected with a `400`.
* **Not combinable with frames.** Reference inputs (`reference_images`, `reference_videos`, `reference_audios`) cannot be combined with the first/last frame inputs (`image`, `last_frame`).
* **Provider scope.** Reference videos and audio are only supported on Seedance 2.x models; sending them to other models returns a `400`. **KLING v3.0** does not support any reference inputs (`reference_images`, `reference_videos`, `reference_audios`); use first/last frame inputs instead.
* **Moderation still applies.** The output is subject to the provider's content moderation. Blocked generations finish as `failed` and are logged with a `content_filter` finish reason. The [gateway content filter](https://docs.vichar.io/resources/error-handling#gateway-content-filter) may also reject the request with a `403` before a job is created.
#### Examples [#examples]
Reference images only (subjects / style):
```bash
curl -X POST "https://api.vichar.io/v1/videos" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-0",
"prompt": "The subject walks through a neon-lit market at night",
"seconds": 5,
"size": "1280x720",
"reference_images": [
{ "image_url": "https://example.com/subject.png" },
{ "image_url": "https://example.com/style.png" }
]
}'
```
Reference video only (motion / scene — let the clip drive the output):
```bash
curl -X POST "https://api.vichar.io/v1/videos" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-0",
"prompt": "adapt this to show more detail",
"seconds": 5,
"size": "1280x720",
"reference_videos": ["https://example.com/reference-motion.mp4"]
}'
```
All three reference types combined:
```bash
curl -X POST "https://api.vichar.io/v1/videos" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-0",
"prompt": "The subject performs the choreography from the reference video",
"seconds": 5,
"size": "1280x720",
"reference_images": [
{ "image_url": "https://example.com/subject.png" }
],
"reference_videos": [
"https://example.com/reference-motion.mp4"
],
"reference_audios": [
"https://example.com/reference-track.mp3"
]
}'
```
### Not supported yet [#not-supported-yet]
* multipart uploads
* `n` values other than `1`
* remix/list/delete video endpoints
## Create a Video [#create-a-video]
Credits-billed video jobs are gated on their estimated cost before anything is submitted upstream. The estimate is the model's per-second rate for the requested resolution (the audio-inclusive rate when the model prices audio separately) times the duration, plus any input-image price. Your organization must hold at least that amount, and never less than `$1.00`, beyond what its still-running video jobs already reserve; otherwise the request fails with `402`. The estimate is reserved when the job is created and settled to the actual cost when the job finishes, so a burst of submissions cannot spend more than the balance covers.
Pricing is per second of generated video. For Seedance and KLING v3.0, enabling audio can increase the per-second rate on models that price audio and video separately.
Veo 3.1:
| Model | Provider | Supported sizes | Price |
| ------------------------------- | --------------- | ------------------------------------------------ | ---------------- |
| `veo-3.1-generate-preview` | `google-vertex` | `1280x720`, `720x1280`, `1920x1080`, `1080x1920` | `$0.40 / second` |
| `veo-3.1-fast-generate-preview` | `google-vertex` | `1280x720`, `720x1280`, `1920x1080`, `1080x1920` | `$0.15 / second` |
| `veo-3.1-generate-preview` | `google-vertex` | `3840x2160`, `2160x3840` | `$0.60 / second` |
| `veo-3.1-fast-generate-preview` | `google-vertex` | `3840x2160`, `2160x3840` | `$0.35 / second` |
Seedance (ByteDance):
| Model | Provider | Resolution | With audio | Video only |
| ------------------- | ----------- | ---------- | ------------------- | ------------------- |
| `seedance-2-5` | `bytedance` | 480p | `$0.1028 / second` | `$0.1028 / second` |
| `seedance-2-5` | `bytedance` | 720p | `$0.2311 / second` | `$0.2311 / second` |
| `seedance-2-5` | `bytedance` | 1080p | `$0.52 / second` | `$0.52 / second` |
| `seedance-2-0` | `bytedance` | 720p | `$0.1512 / second` | `$0.1512 / second` |
| `seedance-2-0` | `bytedance` | 1080p | `$0.3402 / second` | `$0.3402 / second` |
| `seedance-2-0-fast` | `bytedance` | 720p | `$0.121 / second` | `$0.121 / second` |
| `seedance-2-0-mini` | `bytedance` | 480p | `$0.0378 / second` | `$0.0378 / second` |
| `seedance-2-0-mini` | `bytedance` | 720p | `$0.0756 / second` | `$0.0756 / second` |
| `seedance-1-5-pro` | `bytedance` | 720p | `$0.05184 / second` | `$0.02592 / second` |
| `seedance-1-5-pro` | `bytedance` | 1080p | `$0.1166 / second` | `$0.05832 / second` |
KLING (AtlasCloud):
| Model | Provider | Resolution | With audio | Video only |
| ------------------ | ------------ | ---------- | ----------------- | ----------------- |
| `kling-v3-0` | `atlascloud` | 720p | `$0.126 / second` | `$0.084 / second` |
| `kling-v3-0` | `atlascloud` | 1080p | `$0.168 / second` | `$0.112 / second` |
| `kling-v3-0` | `atlascloud` | 4k | `$0.42 / second` | `$0.42 / second` |
| `kling-v3-0-turbo` | `atlascloud` | 720p | `$0.168 / second` | n/a (audio only) |
| `kling-v3-0-turbo` | `atlascloud` | 1080p | `$0.21 / second` | n/a (audio only) |
```bash
curl -X POST "https://api.vichar.io/v1/videos" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "veo-3.1-generate-preview",
"prompt": "A cinematic aerial shot flying above a rainforest waterfall at sunrise",
"seconds": 8,
"size": "1920x1080"
}'
```
Example response:
```json
{
"id": "v_123",
"object": "video",
"model": "veo-3.1-generate-preview",
"status": "queued",
"progress": 0,
"created_at": 1773600000,
"completed_at": null,
"expires_at": null,
"error": null
}
```
## Retrieve Job Status [#retrieve-job-status]
```bash
curl "https://api.vichar.io/v1/videos/v_123" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY"
```
Typical statuses:
* `queued`
* `in_progress`
* `completed`
* `failed`
* `canceled`
* `expired`
Once a job reaches a terminal status and has been billed, the response (and the signed callback payload) includes its cost in USD:
```json
"usage": {
"cost": 0.4,
"cost_details": { "video_output_cost": 0.4, "image_input_cost": 0 }
}
```
Failed, canceled, and expired jobs report a cost of `0`.
`google-vertex` follows Vertex AI's long-running operation flow. The gateway submits Veo generation with `predictLongRunning`, polls with `fetchPredictOperation`, and streams the final bytes through the gateway content endpoint once the operation is done.
`bytedance` uses the ModelArk `/contents/generations/tasks` endpoint. The gateway submits the job, polls the upstream task status, and exposes the final video bytes through the gateway content endpoint once the task succeeds.
`atlascloud` uses the AtlasCloud `/api/v1/model/generateVideo` endpoint and polls `/api/v1/model/prediction/{id}` for status. The gateway resolves the upstream KLING variant (standard, turbo, or 4k) and task type (text-to-video or image-to-video) from your request, then streams the final video bytes through the gateway content endpoint once the prediction completes.
## Download the Video [#download-the-video]
Once the job is complete, stream the resulting video bytes from the content endpoint:
```bash
curl "https://api.vichar.io/v1/videos/v_123/content" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
--output video.mp4
```
## Signed Callbacks [#signed-callbacks]
Vichar can notify your application when the job reaches a terminal state.
```bash
curl -X POST "https://api.vichar.io/v1/videos" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "veo-3.1-fast-generate-preview",
"prompt": "A slow-motion close-up of waves crashing against black volcanic rock",
"seconds": 8,
"callback_url": "https://example.com/webhooks/video",
"callback_secret": "whsec_your_secret_here"
}'
```
### Delivery behavior [#delivery-behavior]
* Callbacks are sent only for terminal states in v1
* Event types are `video.completed` and `video.failed`
* Deliveries retry with exponential backoff on network errors, timeouts, and non-2xx responses
* Each attempt is recorded internally in the webhook delivery log table
### Headers [#headers]
* `webhook-id`
* `webhook-timestamp`
* `webhook-signature`
### Signature format [#signature-format]
Vichar signs the string:
```text
{webhook-id}.{webhook-timestamp}.{raw-request-body}
```
using HMAC-SHA256 with your `callback_secret`, then sends:
```text
webhook-signature: v1,{base64_signature}
```
### Verification example [#verification-example]
```ts
import { createHmac, timingSafeEqual } from "node:crypto";
function verifyWebhook(
body: string,
webhookId: string,
webhookTimestamp: string,
webhookSignature: string,
secret: string,
) {
const expected = createHmac("sha256", secret)
.update(`${webhookId}.${webhookTimestamp}.${body}`)
.digest("base64");
const provided = webhookSignature.replace(/^v1,/, "");
return timingSafeEqual(Buffer.from(expected), Buffer.from(provided));
}
```
## Related Docs [#related-docs]
* [Image Generation](https://docs.vichar.io/features/image-generation)
* [Routing](https://docs.vichar.io/features/routing)
* [Models API](https://docs.vichar.io/v1_models)
### Playback and seeking [#playback-and-seeking]
Video content endpoints forward `Range` and `If-Range` requests to the upstream
storage service and preserve partial-content responses. Inline video results also
support single byte ranges. A satisfiable range returns `206` with `Content-Range`;
an unsatisfiable range returns `416`. This lets native players load and seek
without downloading the entire video first when the upstream supports ranges.
# Vision Support
URL: https://docs.vichar.io/features/vision
Vichar supports vision-enabled models that can analyze and describe images. You can provide images via HTTPS URLs or inline base64-encoded data.
## Vision-Enabled Models [#vision-enabled-models]
You can find all vision-enabled models on our [models page with vision filter](https://app.vichar.io/dashboard). These models can process both text and image content in the same request.
## Image Formats [#image-formats]
### Using HTTPS URLs [#using-https-urls]
You can provide any publicly accessible HTTPS URL pointing to an image:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What do you see in this image?"
},
{
"type": "image_url",
"image_url": {
"url": "https://example.com/image.jpg"
}
}
]
}
]
}'
```
### Using Base64 Inline Data [#using-base64-inline-data]
You can also provide images as base64-encoded data URIs:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Describe this image"
},
{
"type": "image_url",
"image_url": {
"url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQEASABIAAD..."
}
}
]
}
]
}'
```
## Content Array Format [#content-array-format]
When using vision models, the `content` field should be an array containing both text and image content blocks:
* **Text content**: `{"type": "text", "text": "Your message"}`
* **Image content**: `{"type": "image_url", "image_url": {"url": "image_url_or_data_uri"}}`
## Multiple Images [#multiple-images]
You can include multiple images in a single request:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Compare these two images"
},
{
"type": "image_url",
"image_url": {
"url": "https://example.com/image1.jpg"
}
},
{
"type": "image_url",
"image_url": {
"url": "https://example.com/image2.jpg"
}
}
]
}
]
}'
```
## Simple String Content [#simple-string-content]
For vision models, you can still use simple string content for text-only
messages. The array format is only required when including images.
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "Hello! How can you help me today?"
}
]
}'
```
## Supported Image Types [#supported-image-types]
Vision models typically support common image formats including:
* JPEG (.jpg, .jpeg)
* PNG (.png)
* WebP (.webp)
* GIF (.gif)
The specific formats supported may vary by model provider. Check the individual model documentation for format limitations and file size restrictions.
## Error Handling [#error-handling]
If an image URL is inaccessible or the image format is unsupported, the gateway will handle the error gracefully and may substitute a placeholder or error message in the request to the underlying model.
# Native Web Search
URL: https://docs.vichar.io/features/web-search
Vichar supports native web search capabilities that allow models to access real-time information from the internet. This feature is useful for answering questions about current events, recent news, live data, and other time-sensitive information that may not be in the model's training data.
## How It Works [#how-it-works]
When you include the `web_search` tool in your request, the model can search the web to gather relevant information before generating a response:
1. You send a request with the `web_search` tool enabled
2. The model determines if web search is needed based on the query, unless you
[require one](#requiring-a-search)
3. If needed, the model performs web searches to gather current information
4. The model synthesizes the search results and generates a response
5. Citations are included in the response to show information sources
## Supported Providers [#supported-providers]
Native web search is available on select models. See all models with native web search support on our [models page](https://app.vichar.io/dashboard).
**Perplexity Sonar changes on September 25, 2026.** Perplexity retires its
Sonar chat completions API on September 27. On September 25,
`perplexity/sonar` moves to Perplexity's Agent API: the model id, the request
shape and the response fields stay the same, and the flat per-request fee is
replaced by per-search billing. `perplexity/sonar-pro` and
`perplexity/sonar-reasoning-pro` have no equivalent on that API and stop being
routable on September 27. Read the \[full.
## Basic Usage [#basic-usage]
To enable web search, add the `web_search` tool to your request:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.2",
"messages": [
{
"role": "user",
"content": "What is the current weather in San Francisco?"
}
],
"tools": [
{
"type": "web_search"
}
]
}'
```
### Example Response [#example-response]
```json
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1234567890,
"model": "openai/gpt-5.2",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The current weather in San Francisco is 57°F (14°C) with mostly cloudy skies...",
"annotations": [
{
"type": "url_citation",
"url": "https://weather.com/...",
"title": "San Francisco Weather"
}
]
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 15,
"completion_tokens": 150,
"total_tokens": 165,
"cost": 0.0315
}
}
```
## Web Search Options [#web-search-options]
The `web_search` tool accepts optional configuration parameters:
### User Location [#user-location]
Provide location context to get more relevant local search results:
```json
{
"type": "web_search",
"user_location": {
"city": "San Francisco",
"region": "California",
"country": "US",
"timezone": "America/Los_Angeles"
}
}
```
### Search Context Size [#search-context-size]
Control the amount of web content retrieved (OpenAI only):
```json
{
"type": "web_search",
"search_context_size": "medium"
}
```
Available values:
* `low` - Minimal search context, faster responses
* `medium` - Balanced context (default)
* `high` - Maximum search context, more comprehensive
### Max Uses [#max-uses]
Limit the number of searches per request (provider-dependent):
```json
{
"type": "web_search",
"max_uses": 3
}
```
### Domain Filters [#domain-filters]
Restrict which domains the model may search with `allowed_domains`, or exclude domains with `blocked_domains`. Both fields are accepted independently; provider support varies — currently the filters reach Anthropic-served requests, and other providers ignore them. Anthropic accepts only one of the two, so when both are set the gateway forwards `allowed_domains` and drops `blocked_domains`:
```json
{
"type": "web_search",
"allowed_domains": ["example.com"]
}
```
### Shorthand: `web_search: true` [#shorthand-web_search-true]
As a shortcut, set the top-level `web_search` body field to `true` instead of adding the tool — the gateway injects a default `web_search` tool for you (it has no effect if the tool is already present):
```json
{
"model": "gpt-5.2",
"messages": [{ "role": "user", "content": "What happened in tech today?" }],
"web_search": true
}
```
## Requiring a Search [#requiring-a-search]
By default the `web_search` tool offers the model a search and lets it judge
whether the question needs one. To require a search on every request, set
`tool_choice`:
```json
{
"model": "...",
"messages": [{ "role": "user", "content": "What shipped in AI this week?" }],
"tools": [{ "type": "web_search" }],
"tool_choice": { "type": "web_search" }
}
```
Reach for this sparingly. A forced search is billed on every request that
carries it, and the retrieved snippets are appended to your prompt, so they are
billed as input tokens too — on a question the model could have answered from
memory, you pay for both and gain nothing. If you are building a chat interface
with a "web search" toggle, leaving the tool attached with the default
`tool_choice` is usually what you want, so that follow-ups like "shorter,
please" do not trigger a search.
A few upstreams have no model-elected search at all and can only search when
asked to. Requiring a search is the only way to use their search; without it
they behave like a model that decided not to search, and the gateway prefers to
route a merely offered tool to a provider that can make that decision for
itself. You can find the models with native web search on the
[models page](https://app.vichar.io/dashboard).
## Using with SDKs [#using-with-sdks]
### OpenAI SDK (Python) [#openai-sdk-python]
```python
from openai import OpenAI
client = OpenAI(
base_url="https://api.vichar.io/v1",
api_key="your-api-key"
)
response = client.chat.completions.create(
model="gpt-5.2",
messages=[
{"role": "user", "content": "What are the latest news headlines today?"}
],
tools=[{"type": "web_search"}]
)
print(response.choices[0].message.content)
```
### OpenAI SDK (TypeScript) [#openai-sdk-typescript]
```typescript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: "your-api-key",
});
const response = await client.chat.completions.create({
model: "gpt-5.2",
messages: [{ role: "user", content: "What are the latest tech news?" }],
tools: [{ type: "web_search" }],
});
console.log(response.choices[0].message.content);
```
## Streaming [#streaming]
Web search works with streaming responses. Citations are included in the final chunks:
```bash
curl -X POST "https://api.vichar.io/v1/chat/completions" \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.2",
"messages": [
{"role": "user", "content": "What is the current stock price of Apple?"}
],
"tools": [{"type": "web_search"}],
"stream": true
}'
```
## Citations and Sources [#citations-and-sources]
Web search responses include citations to show where information was sourced from. These appear in the `annotations` field of the message:
```json
{
"annotations": [
{
"type": "url_citation",
"url": "https://example.com/article",
"title": "Article Title",
"start_index": 0,
"end_index": 50,
"date": "2026-09-16",
"last_updated": "2026-09-17"
}
]
}
```
Citation format may vary slightly between providers, but Vichar normalizes
them into a consistent structure. `date` and `last_updated` are only present
when the provider reports them for that source.
Search-grounded models that return their sources as a set rather than as inline
citations also include a top-level `search_results` array alongside `choices`,
with one entry per source (`url`, `title`, `snippet`, `date`, `last_updated`).
Streaming responses carry it on the chunk that delivers the sources.
## Cost Tracking [#cost-tracking]
Web search costs are rolled into the total `cost` reported in the usage object:
```json
{
"usage": {
"prompt_tokens": 15,
"completion_tokens": 150,
"total_tokens": 165,
"cost": 0.0125,
"cost_details": {
"upstream_inference_cost": 0.0115,
"upstream_inference_prompt_cost": 0.0015,
"upstream_inference_completions_cost": 0.01,
"total_cost": 0.0125,
"input_cost": 0.0015,
"output_cost": 0.01,
"web_search_cost": 0.001
}
}
}
```
Web search is billed at $0.01 per search call for reasoning models (GPT-5, o-series) and $0.025 per call for non-reasoning models. The web search charge is included in the top-level `cost` value and surfaced separately as `cost_details.web_search_cost`.
## Combining with Function Tools [#combining-with-function-tools]
You can use web search alongside regular function tools:
```json
{
"tools": [
{ "type": "web_search" },
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": { "type": "string" }
}
}
}
}
]
}
```
Some dedicated search models only support web search and do not support
additional function tools. Use `gpt-5.2` or other GPT-5 series models if you
need both web search and function tools.
## Use Cases [#use-cases]
### Current Events and News [#current-events-and-news]
```json
{
"messages": [
{ "role": "user", "content": "What are the major news stories today?" }
],
"tools": [{ "type": "web_search" }]
}
```
### Real-Time Data [#real-time-data]
```json
{
"messages": [
{ "role": "user", "content": "What is the current price of Bitcoin?" }
],
"tools": [{ "type": "web_search" }]
}
```
### Research and Fact-Checking [#research-and-fact-checking]
```json
{
"messages": [
{
"role": "user",
"content": "What are the latest findings on climate change?"
}
],
"tools": [{ "type": "web_search" }]
}
```
### Local Information [#local-information]
```json
{
"messages": [
{
"role": "user",
"content": "What restaurants are open near me right now?"
}
],
"tools": [
{
"type": "web_search",
"user_location": {
"city": "New York",
"country": "US"
}
}
]
}
```
## Best Practices [#best-practices]
1. **Use GPT-5.2**: For the best web search experience with full tool support, use `gpt-5.2`
2. **Provide location context**: When queries are location-dependent, include `user_location` for more relevant results
3. **Monitor costs**: Web search incurs per-query costs in addition to token costs
4. **Check citations**: Always review the citations in responses to verify information sources
5. **Use streaming**: For user-facing applications, enable streaming to show responses as they're generated
## Error Handling [#error-handling]
If you try to use web search with a model that doesn't support it:
```json
{
"error": {
"message": "Model gpt-4o does not support native web search. Remove the web_search tool or use a model that supports it. See https://app.vichar.io/dashboard for supported models.",
"type": "invalid_request_error"
}
}
```
To avoid this error, only use the `web_search` tool with [native web search enabled models](https://app.vichar.io/dashboard).
# Activity
URL: https://docs.vichar.io/learn/activity
The Activity page shows a real-time log of every API request routed through Vichar. Use it to debug requests, monitor performance, and track costs per call.
## Filters [#filters]
Filter the activity log using the controls at the top:
| Filter | Description |
| --------------------------- | --------------------------------------------------------------------------- |
| **Time range** | Filter by a specific time period |
| **Errors** | Show only errored requests, or narrow to client, gateway or upstream errors |
| **Unified reasons** | Filter by completion reason (e.g., stop, length, error) |
| **Providers** | Show requests for specific providers only |
| **Models** | Show requests for specific models only |
| **Custom header key/value** | Filter by custom metadata headers attached to requests |
## Activity List [#activity-list]
Each activity entry shows:
* **Status icon** — Green checkmark for completed, red circle for errors
* **Response preview** — First line of the model's response (when available)
* **Model** — The provider and model used (e.g., `google-vertex/gemini-3-pro-image`)
* **Cache status** — Whether the response was served from cache
* **Tokens** — Total tokens consumed (input + output)
* **Duration** — How long the request took
* **Cost** — Inference cost for the request
* **Source** — Where the request originated from
* **Discount** — Any discount applied (e.g., "20% off")
* **Status badge** — `completed`, `upstream_error`, `gateway_error`, etc.
* **Timestamp** — Relative time (e.g., "about 4 hours ago")
### Actions per Entry [#actions-per-entry]
* **Open in new tab** — View the full request detail in a new browser tab
* **Expand** — Expand inline to see more details
## Activity Detail [#activity-detail]
Click on any activity entry to view its full detail page.
### Summary Cards [#summary-cards]
Five cards at the top provide a quick overview:
| Card | Description |
| ------------------ | ------------------------------- |
| **Duration** | Total request time in seconds |
| **Tokens** | Total tokens consumed |
| **Throughput** | Tokens per second |
| **Inference Cost** | Cost charged for this request |
| **Cache** | Whether the response was cached |
### Request Section [#request-section]
Details about the original request:
* **Requested Model** — The model ID sent in the API call
* **Used Model** — The actual model that served the request
* **Model Mapping** — The underlying model identifier
* **Provider** — The provider that handled the request
* **Requested Provider** — The provider specified in the request
* **Streamed** — Whether the response was streamed
* **Canceled** — Whether the request was canceled
* **Source** — The application or service that made the request
### Tokens Section [#tokens-section]
A detailed token breakdown:
* Prompt Tokens, Completion Tokens, Total Tokens
* Reasoning Tokens (for reasoning models)
* Image Input/Output Tokens (for vision/image models)
* Response Size
### Routing Section [#routing-section]
How Vichar routed the request:
* **Selection** — The routing strategy used (e.g., `direct-provider-specified`)
* **Available** — Providers that were available for this model
* **Provider Scores** — Scoring breakdown showing availability, uptime, and latency for each provider
### Parameters Section [#parameters-section]
The model parameters sent with the request:
* Temperature, Max Tokens, Top P
* Frequency Penalty, Reasoning Effort
* Response Format
# Analytics
URL: https://docs.vichar.io/learn/analytics
The Analytics page shows where a project's spend actually goes. It leads with the project's totals for the range, then breaks usage down — as a ranking and over time — so you can see what drives cost, requests, and tokens across any date range.
Open it from the **Analytics** item in the project sidebar. Like the rest of the dashboard, the page respects the shared date-range picker at the top, so every chart reflects the same window.
## Totals [#totals]
Three cards at the top show the project's **total cost**, **requests**, and **tokens** for the selected range, each with a sparkline of its daily values. They are computed from the same data as the charts below, so the total and the breakdown always agree.
Request and token counts use commas every three digits, regardless of browser locale. Chart axes and compact summaries use `k`, `M`, `B`, or `T` for thousands through trillions; hover chart points to inspect full counts.
## Choosing a breakdown [#choosing-a-breakdown]
The **Group by** selector switches which dimension the two charts break the project down by:
| Breakdown | What it shows |
| ----------- | ---------------------------------------------- |
| **Model** | Spend per model (the default) |
| **API key** | Spend per API key in the project |
| **User** | Spend per team member — owners and admins only |
**Breakdown by user** answers "what is each person on this team costing us?" in the same view as the project total — see [Member Analytics](https://docs.vichar.io/learn/member-analytics).
Spend is attributed to whoever **created** the API key that made each request. Traffic from platform keys has no member behind it, so it is shown as **Unattributed** rather than dropped, keeping the per-member rows summing to the project total.
## Cost by Model [#cost-by-model]
A horizontal bar chart that ranks the models used in the selected range. Switch the metric with the tabs above the chart:
| Tab | What it ranks |
| ------------ | -------------------------------------- |
| **Cost** | Total spend per model, in USD |
| **Requests** | Number of requests routed to the model |
| **Tokens** | Total tokens (input + output) |
Use it to spot the one or two models responsible for most of your bill, or to confirm that traffic is spread the way you expect.
## Cost by Model Over Time [#cost-by-model-over-time]
Choose **Line** to compare individual series or **Bar** to stack their values across the date range. Your chart-style choice is remembered across visits and also applies to the API key, user, and token charts. It has the same **Cost / Requests / Tokens** tabs, plus a **Mappings / Canonical** toggle:
* **Mappings** — each model variant is shown separately, exactly as it was requested (for example a custom provider mapping is kept distinct from the built-in model).
* **Canonical** — variants of the same underlying model are collapsed into one canonical model (the provider prefix and tag are dropped), so `openai/gpt-5.5` and a custom mapping of it count as a single line.
Switch to **Canonical** when you care about the underlying model's total footprint; stay on **Mappings** when you need to compare specific routes.
The **Mappings / Canonical** toggle is specific to models — the API key and user breakdowns rank named entities, so there is nothing to collapse.
## Token breakdown [#token-breakdown]
**Tokens over time** separates uncached input, cache reads, and output. Input includes cache-write tokens; the note below the chart shows how many were written to cache. Cache reads appear only in the cache series, so the three categories do not double-count input tokens.
The model selector above the chart narrows all three series to a single model, so you can see where one model's cache reads or output actually land. It lists the models used in the selected range and is available on the model breakdown.
The billing-mode selector filters spend and requests; historical token counters are combined.
For budget warnings, upcoming model retirements, and provider issues, enable [notifications](https://docs.vichar.io/features/notifications) from the bell in the dashboard header.
## How the data is computed [#how-the-data-is-computed]
These charts derive from the same activity data the rest of the dashboard already reads — there is no separate analytics pipeline to wait on. Aggregation happens per time bucket and is timezone-correct, so totals line up with the Activity and Usage pages for the same range.
For a per-key view of the same breakdowns, see [API Keys](https://docs.vichar.io/learn/api-keys#per-key-statistics). For an org-wide, per-person view, see [Member Analytics](https://docs.vichar.io/learn/member-analytics).
# API Keys
URL: https://docs.vichar.io/learn/api-keys
The API Keys page is the main place to create, secure, and operate the keys
your apps use to authenticate with Vichar.
Use this page to:
* Create project-specific API keys
* Set all-time and recurring spend limits per key
* Set an expiration (TTL) so a key disables itself automatically
* Track usage for each key, including the active recurring window
* Rename keys and enable or disable them without deleting them
* Roll a key's secret in place if it may have leaked
* Configure IAM rules for model, provider, and pricing access
API keys are shown in full only once, immediately after creation. Copy and
store them securely before closing the dialog. New and rolled secrets are
stored only as keyed HMAC-SHA-256 fingerprints; authentication fingerprints
the presented secret and compares the result.
## Creating an API Key [#creating-an-api-key]
Click **Create API Key** and configure:
* **Name**: A label such as `production`, `staging`, or `ci`
* **Expiration (TTL)**: An optional time-to-live after which the key disables itself
* **All-time usage limit**: An optional lifetime spend cap for the key
* **Recurring usage limit**: An optional spend cap that resets on a schedule
Recurring limits support:
* Minimum window: **1 hour**
* Maximum window: **12 months**
* Units: **hour**, **day**, **week**, or **month**
This is useful when you want a key to stay below a fixed budget per hour, day,
week, or month, while still keeping a separate lifetime cap if needed.
## Expiration (TTL) [#expiration-ttl]
Turn on **Set expiration (TTL)** when creating a key to give it a limited
lifetime. Choose a value and a unit — **minutes**, **hours**, or **days** — and
the key is disabled automatically once that time passes. Leave it off for a key
that never expires.
Expired keys show an **Expired** indicator in the list and move to the
**Inactive** tab. To use one again, reactivate it and pick a **new future
expiration**:
* **Activate** an expired key and you'll be prompted to set a fresh TTL before it
comes back online
* Keys with no TTL, or whose TTL is still in the future, can be enabled and
disabled without setting a new expiration
This makes TTL keys ideal for temporary access — short-lived demos, CI runs, or
contractor keys that should not linger.
## Usage Limits [#usage-limits]
Each API key can enforce two independent limit types:
| Limit Type | What it does |
| ------------------------- | --------------------------------------------------------------- |
| **All-time usage limit** | Stops the key after it reaches a lifetime spend threshold |
| **Recurring usage limit** | Stops the key after it reaches the budget for the active window |
Examples:
* `$50` all-time for a temporary integration key
* `$10 / 1 day` for a development key
* `$500 / 1 month` for a production service key
If a key hits either limit, requests using that key are rejected until the key
is updated or, for recurring limits, the next window begins.
### How recurring windows work [#how-recurring-windows-work]
Recurring usage is tracked separately from total lifetime usage.
* The dashboard shows the key's **Current Period** usage
* The active window also shows when it **resets**
* When the configured window expires, usage for that window resets automatically
* Updating the recurring limit configuration resets the current window and starts
a new one
Usage reflects requests billed to your Vichar credits.
## API Keys List [#api-keys-list]
Each key in the list shows:
| Field | Description |
| ------------------ | ------------------------------------------------------------- |
| **Name** | The label you assigned to the key |
| **API Key** | A masked preview of the key |
| **Status** | Whether the key is active or inactive, plus its expiry if set |
| **Created** | When the key was created |
| **Usage** | Total tracked usage for the key |
| **Current Period** | Spend in the active recurring window, if configured |
| **Limits** | All-time and recurring limit summary |
| **IAM Rules** | Whether model/provider/pricing access controls are configured |
Every configured limit renders as a gauge, so a key's headroom is visible without
opening it. A key is flagged **approaching** once usage reaches 80% of a limit and
**reached** once usage meets it — the all-time limit stays blocked until it is
raised or removed, while a recurring limit clears itself on its reset date.
Three filters sit above the list:
* **Status** — active (default), inactive, or all
* **Creator** — keys you created, or everyone's (hidden for developers, who only
ever see their own)
* **Limits** — narrow to just the keys approaching or at a limit
A filter that would leave the list empty resets itself, so you never land on a
blank table.
## Actions [#actions]
The actions menu on each key offers:
* **View Statistics**: Open a dedicated analytics page for that key (see below)
* **Manage IAM Rules**: Restrict which models, providers, or pricing tiers the key can use
* **Rename Key**: Change the key's label without touching its secret or history
* **Activate / Deactivate Key**: Pause usage without deleting the key (reactivating
an expired key prompts for a new expiration)
* **Roll Key**: Generate a new secret for the key and invalidate the old one (see below)
* **Update limits**: Change all-time or recurring limits
* **Delete**: Permanently remove the key
The Playground key appears with a **Managed** badge and only exposes its
statistics. It cannot be renamed, rolled, deactivated, limited, assigned IAM
rules, or deleted. Members with the developer role can only modify keys they
created themselves.
## Rolling a Key [#rolling-a-key]
**Roll Key** replaces a key's secret while keeping the key itself. The name,
usage history and statistics, all-time and recurring limits (including the
current period window), IAM rules, and expiration all stay exactly as they were —
only the secret changes.
This is the fastest fix when a secret leaks into a commit, a log, or a shared
environment: you cut off the exposed value without recreating the key or losing
its spend tracking.
1. Open the key's actions menu and choose **Roll Key**
2. Confirm in the dialog
3. Copy the new secret and update every client using the old one
The previous secret stops working immediately, and the new one is shown only
once. Requests still using the old secret fail with `401 Unauthorized`.
## Per-Key Statistics [#per-key-statistics]
The **View Statistics** action opens a dedicated page scoped to a single API
key, so you can see exactly what that key is doing without filtering the whole
project.
The page respects the shared date-range picker and shows:
* **Summary cards** — the key's cost, tokens, requests, and error rate for the
selected range.
* **Cost by Model** — a horizontal bar chart ranking the key's models by cost,
requests, or tokens.
* **Cost by Model Over Time** — a stacked area chart of the same metrics, with a
Mappings / Canonical toggle.
These are the same breakdowns as the project [Analytics](https://docs.vichar.io/learn/analytics) page,
narrowed to the one key — useful for confirming a key is healthy and spending on
the models you expect.
## IAM Rules [#iam-rules]
IAM rules let you narrow what an API key is allowed to access.
Supported rule types include:
* **Allow/Deny models**
* **Allow/Deny providers**
* **Allow/Deny pricing**
* **Allow/Deny IP ranges (CIDR)** — on request (contact us)
Use IAM rules when you want a key to be valid, but only for a specific subset of
models or providers. For a deeper explanation, see the [API Keys & IAM Rules
feature page](https://docs.vichar.io/features/api-keys).
Org admins can additionally set [member-level IAM
rules](https://docs.vichar.io/features/api-keys#member-level-iam-rules) on the [Team
page](https://docs.vichar.io/learn/team). Those act as a ceiling for every key the member creates:
key rules can only narrow access within them, never expand it. The IAM page
shows a notice when organization-level restrictions apply to your keys.
## Plan Limits [#plan-limits]
The page also shows how many API keys you are using relative to your plan
allowance. The cap counts **active** keys across every project in the
organization, so deactivated and deleted keys do not count against it.
* **Free**: Standard API key count limit
* **Pro**: A larger allowance
* **Enterprise**: Custom limits
Owners and admins can also set a per-member key cap on the [Team
page](https://docs.vichar.io/learn/team), which applies on top of the organization-wide allowance.
If you reach the limit, the **Create API Key** button is disabled until you
delete or deactivate unused keys, or upgrade.
# Billing
URL: https://docs.vichar.io/learn/billing
The Billing page is your central hub for managing credits, payment methods, and automatic top-ups.
## Credits [#credits]
Displays your current credit balance. Credits are consumed as you make API requests through the gateway and never expire. Click **Top Up Credits** to add more credits to your account.
Payments are processed by Dodo Payments, our merchant of record — checkout happens on their secure hosted page.
## Fees [#fees]
Top-ups are charged the credit amount plus a flat 5% platform fee. Any applicable tax is calculated and added at checkout by Dodo Payments. The minimum top-up is $20 and the maximum is $5,000.
The breakdown (credits, platform fee, and total before tax) is shown in the top-up dialog before you continue to checkout, so the charge is always transparent.
## Payment Methods [#payment-methods]
Click **Manage payment methods & invoices** to open the secure Dodo billing portal, where you can update your payment methods and download payment receipts and invoices.
## Auto Top-up Settings [#auto-top-up-settings]
Configure automatic credit top-ups so you never run out:
* **Enable auto top-up** saves a payment method via a secure Dodo checkout and stores it as an auto top-up mandate
* **Threshold** — The credit balance that triggers a top-up
* **Amount** — How many credits to add when the threshold is reached
* **Disable** cancels the mandate and stops automatic purchases
This ensures uninterrupted service by automatically replenishing your credits when they run low.
# Dashboard
URL: https://docs.vichar.io/learn/dashboard
The Dashboard is the first page you see after logging in. It provides a high-level overview of your project's LLM usage, costs, and performance at a glance.
## Date Range [#date-range]
At the top of the page, you can toggle the date range for all dashboard metrics:
* **7 days** — Last 7 days of data (default)
* **30 days** — Last 30 days of data
* **Custom** — Pick a custom start and end date
### Display time zone [#display-time-zone]
Timestamps and chart buckets follow the display time zone set under **Settings →
Account**: either your local zone or UTC. The preference drives both how a
datetime renders and the bucketing the analytics endpoints do, so a chart's bars
and its axis labels always agree. It is stored in a browser cookie rather than
your profile, so the server renders the correct zone on first paint.
## Stat Cards [#stat-cards]
The dashboard displays metric cards in two rows. Organization credits and top-up actions are visible only to organization owners and admins.
Across the dashboard, full request and token counts use commas every three digits. Compact summaries and chart axes use `k`, `M`, `B`, or `T` for thousands through trillions, independent of browser locale.
### Top Row [#top-row]
| Card | Description |
| ------------------------ | ------------------------------------------------------------------------ |
| **Organization Credits** | Your current available credit balance |
| **Total Requests** | Number of API requests in the selected period, with cache hit percentage |
| **Total Cost** | Total inference cost for the period, including storage costs |
| **Total Savings** | Savings from discounts during the selected period |
### Bottom Row [#bottom-row]
| Card | Description |
| ------------------------ | ------------------------------------------------------------------- |
| **Input Tokens & Cost** | Total prompt tokens sent and their associated cost |
| **Output Tokens & Cost** | Total completion tokens received and their associated cost |
| **Cached Tokens & Cost** | Tokens served from cache (if caching is enabled) and the cost saved |
| **Most Used Model** | The model with the highest request count, along with its provider |
## Usage Overview Chart [#usage-overview-chart]
Below the stat cards, a chart visualizes your usage over time:
* **Costs** — Choose **Total** for daily cost bars or **Breakdown** to split input, output, and cached input costs.
* **Requests** — Shows request volume over time.
The chart is filtered by the currently selected project.
### Compare periods [#compare-periods]
Choose **Compare** to compare another date range by elapsed day:
* **Previous period** uses the same number of days immediately before the active range.
* **Week over week** uses seven days starting on the date you choose.
* **Month over month** uses one calendar month starting on the date you choose. August 1 runs through August 31; February 1 runs through February 28, or February 29 in a leap year.
* **Custom range** lets you choose both comparison dates.
Week and month comparisons always end before the active range begins. Their duration is independent of the main date filter, so a month comparison still shows the full month when the dashboard displays seven days. The chart includes every day in the longer range and leaves the shorter series blank after it ends. Comparison is unavailable for all-time ranges.
## Quick Actions [#quick-actions]
A sidebar panel provides shortcuts to common tasks:
* **Manage API Keys** — Go to the API Keys page
* **View Activity** — See detailed request logs
* **Usage & Metrics** — Dive into usage analytics
* **Model Usage** — View per-model usage breakdown
## Cost Breakdown [#cost-breakdown]
A donut chart showing how your costs are distributed across different models and providers. Each segment is color-coded and labeled with the model name and cost, making it easy to identify your biggest cost drivers.
## Errors & Reliability [#errors--reliability]
Displays two key reliability metrics:
* **Error Rate** — Percentage of failed requests over the selected period
* **Uptime** — Gateway availability percentage
## Recent Activity [#recent-activity]
A table showing your most recent API requests with key details like model, status, tokens, duration, and cost. Click any entry to view the full request detail.
## Header Actions [#header-actions]
The top-right corner provides these actions:
* **Create API Key** — Quickly create a new API key for your project
* **Top Up Credits** — Add credits to your organization balance (organization owners and admins)
# Introduction
URL: https://docs.vichar.io/learn
The Vichar dashboard gives you full control over your LLM API usage, costs, and configuration. This section walks you through every Vichar dashboard page.
Use the organization selector at the top of the dashboard sidebar to switch organizations.
## AI Gateway [#ai-gateway]
Configure how requests flow through the gateway:
* [**API Keys**](https://docs.vichar.io/learn/api-keys) — Create and manage your API keys, plus per-key statistics
* [**Models**](https://docs.vichar.io/learn/models) — Browse the full model catalog
* [**Structured Outputs**](https://docs.vichar.io/learn/structured-outputs) — Soft JSON output versus strict JSON schema enforcement
* [**Preferences**](https://docs.vichar.io/learn/preferences) — Project-level settings like caching
## Observability [#observability]
Monitor usage, spend, errors, and performance across every request:
* [**Dashboard**](https://docs.vichar.io/learn/dashboard) — Overview of your usage, costs, and performance
* [**Activity**](https://docs.vichar.io/learn/activity) — Detailed logs of every API request
* [**Model Usage**](https://docs.vichar.io/learn/model-usage) — Usage breakdown by model
* [**Usage & Metrics**](https://docs.vichar.io/learn/usage-metrics) — Requests, errors, cache rates, and cost trends
* [**Analytics**](https://docs.vichar.io/learn/analytics) — Cost, requests, and tokens broken down by model
* [**Member Analytics**](https://docs.vichar.io/learn/member-analytics) — Per-member cost and usage breakdowns
## Account & Billing [#account--billing]
Manage your organization, team, and payments:
* [**Billing**](https://docs.vichar.io/learn/billing) — Credits, plans, and payment methods
* [**Transactions**](https://docs.vichar.io/learn/transactions) — Payment and credit history
* [**Invoices**](https://docs.vichar.io/learn/invoices) — Download invoices and credit notes
* [**Refunds**](https://docs.vichar.io/learn/refunds) — Self-service refunds for top-ups, plans, and Reset Passes
* [**Referrals**](https://docs.vichar.io/learn/referrals) — Earn credits by referring others
* [**Org Preferences**](https://docs.vichar.io/learn/org-preferences) — Organization name and billing details
* [**Team**](https://docs.vichar.io/learn/team) — Manage members, roles, and shared developer policy
# Invoices
URL: https://docs.vichar.io/learn/invoices
Every completed payment gets a PDF invoice, and every refund a credit note. The
invoice is emailed to your billing address when the payment succeeds, and both
can be downloaded at any time from [**Transactions**](https://docs.vichar.io/learn/transactions) in
your organization settings.
## Downloading an invoice [#downloading-an-invoice]
1. Open **Settings → Transactions** in the dashboard.
2. Find the payment in the list.
3. Click **Invoice** on that row — the PDF downloads as
`invoice-.pdf`.
Refunds produce a **Credit note** instead of an invoice, downloaded the same
way. It shows the original purchase amount, the percentage refunded, and the
refunded total as a negative line.
Only **completed** payments with a positive amount have a document. Plan
cancellations, plan endings, and free credit gifts have nothing to invoice, so
those rows have no download button.
The invoice number is the transaction ID, so it matches the row you downloaded
it from.
## Putting company details on the invoice [#putting-company-details-on-the-invoice]
Company name, address, tax ID and invoice notes come from your organization's
billing information, under **Settings → Billing**:
| Field | Appears on the invoice as |
| ----------------------- | ---------------------------------------------- |
| **Email Address** | The billing email address invoices are sent to |
| **Company Name** | The billed company |
| **Billing Address** | The billing address block |
| **Tax ID / VAT Number** | Your tax identification number |
| **Invoice Notes** | A free-text footer (e.g. a PO number) |
Fill these in **before** you pay. The emailed invoice is generated at payment
time and keeps the details as they were then, while a download from
Transactions is rendered fresh from your current details — so the two copies
of one invoice number can differ after you edit this page. Treat the emailed
PDF as the record of what was invoiced, and for a lasting change get the
details right before the next payment.
## Related pages [#related-pages]
* [**Transactions**](https://docs.vichar.io/learn/transactions) — the full payment and credit history
* [**Billing**](https://docs.vichar.io/learn/billing) — credits, plans, and payment methods
# Member Analytics
URL: https://docs.vichar.io/learn/member-analytics
Member Analytics breaks your organization's usage down by person, so you can see who is spending what across every project. It lives on the **Team** page; contact us to enable it for your organization.
Member analytics require an organization **owner** or **admin** role, and
members without admin access see an access notice instead of the data.
## Members Table [#members-table]
The Team page adds a usage table sorted by spend, so the heaviest users surface first. Each row shows that member's totals for the selected date range:
| Column | Description |
| -------------- | -------------------------------------------- |
| **Member** | Name and email of the team member |
| **Cost** | Total spend attributed to the member, in USD |
| **Tokens** | Total tokens (input + output) |
| **Requests** | Number of requests |
| **Error rate** | Share of the member's requests that failed |
| **API keys** | How many API keys the member created |
Usage is attributed by **who created each API key** — spend lands on the member who owns the key that made the request, which is the only link between usage and a user.
## Member Detail [#member-detail]
Click a member to open their detail page, scoped to the same date range:
The detail view includes:
* **Summary cards** — the member's cost, tokens, requests, and error rate for the range.
* **Most used** — their top model, provider, and app.
* **Cost by model** — the same breakdown as the project [Analytics](https://docs.vichar.io/learn/analytics) page, scoped to this member.
* **Top providers and top apps** — tables ranking where the member's traffic goes.
This makes it easy to attribute cost to a team, investigate a spike, or confirm a member is using the models and providers you expect.
# Model Usage
URL: https://docs.vichar.io/learn/model-usage
The Model Usage page shows how your API requests are distributed across different LLM models over time.
## Filters [#filters]
Two filters let you narrow down the data:
* **API Key** — Select a specific API key or view usage across all keys
* **Date range** — Choose a time period to analyze
## Usage Chart [#usage-chart]
The main chart displays a time-series breakdown of requests per model. Each model is represented by a different color, making it easy to see:
* Which models are used most frequently
* How usage patterns change over time
* Whether usage is concentrated on a single model or spread across many
This page is useful for understanding your model distribution and identifying opportunities to optimize costs by switching to more cost-effective models for certain workloads.
# Models
URL: https://docs.vichar.io/learn/models
The Models page shows the full Vichar catalog in one directory, with one row per provider mapping.
Every organization member can browse this page, including project-scoped `developer` members, so your whole team can see which providers and models are available without needing admin access.
## The directory [#the-directory]
The directory is the same table you know from the public [models page](https://app.vichar.io/dashboard), with one row per provider mapping:
| Column | Description |
| ---------------------- | ---------------------------------------------------------------------- |
| **Provider** | The provider serving the model |
| **Model ID** | The model id; the copy button copies the exact string to request |
| **Input / Output $/M** | Prices per million tokens (per-request/per-character where applicable) |
| **Cache Read $/M** | Cached input price where supported |
| **Features** | Capability icons: streaming, vision, tools, reasoning, JSON output, … |
Search matches model names, ids, aliases, and providers. The **Filters** panel offers the same controls as the public directory — use case, capabilities, provider, price and context ranges.
### Lifecycle status [#lifecycle-status]
Every provider mapping carries exactly one lifecycle status, shown as a badge on
the mapping and selectable from the **Status** chips in the filter panel:
| Status | Meaning |
| --------------- | ------------------------------------------------------------------------ |
| **Deprecated** | Still routes; the provider has announced a sunset, so migrate soon |
| **Scheduled** | Still routes, but deactivation is set for a date inside the next 90 days |
| **Deactivated** | No longer routes; requests return errors |
Status resolves by urgency — deactivated beats scheduled beats deprecated — so a
mapping never carries two, and a deactivation further out than the 90-day notice
window stays plain active rather than warning early. The chips are single-select
and write to `?status=`, so a filtered view can be shared.
With no chip selected, the directory lists everything routable and hides
mappings already past their deprecation or deactivation date.
### Time-based pricing [#time-based-pricing]
A few provider mappings bill different rates depending on when the request
arrives. Those cards carry a **Time-based pricing** toggle: switch between
**Peak** and **Off-peak** to see each set of input, cached-input, and output
prices, with the schedule that decides which one applies spelled out underneath —
the peak hours, their time zone, and any days that are off-peak all day.
Billing follows the same schedule automatically at request time; there is
nothing to set on your side. To pay off-peak rates, move batch or backfill work
into the off-peak window.
## Provider performance [#provider-performance]
Open a model's uptime page from the public [models directory](https://app.vichar.io/dashboard) to compare provider requests, errors, latency, and token volume over the last four hours. Chart axes abbreviate large values as thousands (`k`), millions (`M`), billions (`B`), or trillions (`T`); hover a point for the full value with commas separating groups of three digits.
# Org Preferences
URL: https://docs.vichar.io/learn/org-preferences
The Org Preferences page contains settings for your organization's identity and billing information.
## Organization Name [#organization-name]
Update your organization's display name. This name appears throughout the dashboard and in billing communications.
## Billing Email [#billing-email]
Set or update the email address used for billing-related communications, including receipts, invoices, and payment notifications.
## Billing Information [#billing-information]
Configure your organization's billing details for invoices:
| Field | Description |
| ---------------------------------- | ------------------------------------------------------------------------ |
| **Email Address** | Primary email for billing communications |
| **Company Name** (optional) | Your company or organization name for invoices |
| **Billing Address** | Street address, city, state/province, ZIP code, and country |
| **Tax ID / VAT Number** (optional) | Your tax identification or VAT number for tax-compliant invoices |
| **Invoice Notes** (optional) | Custom notes to include on invoices (e.g., PO numbers, department codes) |
## Notification Channels [#notification-channels]
Connect a Slack incoming webhook to post organization-wide alerts, to a Slack channel. Owners and admins can save, test, or replace the webhook. It is stored encrypted and shown masked.
# Preferences
URL: https://docs.vichar.io/learn/preferences
The Preferences page contains project-level settings that control how your project behaves.
## Project Name [#project-name]
Update the display name for your project. This name appears in the sidebar and throughout the dashboard.
## Project Mode [#project-mode]
Configure how your organization handles projects. This setting determines the routing and isolation behavior for API requests within the project.
## Caching [#caching]
Enable or configure response caching for API requests. When enabled, identical requests will return cached responses instead of making new calls to the provider, saving both time and cost.
Response caching must be disabled for every project before the organization can enable zero data retention. While ZDR is active, response caching cannot be enabled and provider prompt-cache markers are stripped.
Learn more about caching in the [Caching feature docs](https://docs.vichar.io/features/caching).
## Danger Zone [#danger-zone]
The Danger Zone section contains irreversible actions:
* **Archive Project** — Permanently archive the project. This action cannot be undone. Archived projects stop processing requests and their API keys become inactive.
# Referrals
URL: https://docs.vichar.io/learn/referrals
The Referrals page lets you earn credits by inviting others to use Vichar.
## Eligibility [#eligibility]
To unlock the referral program, your organization must have at least **$100 in total credit top-ups**. Before reaching this threshold, the page shows:
* A progress bar showing your progress toward $100
* The remaining amount needed to unlock
* An explanation of the 1% earnings model
## Referral Dashboard [#referral-dashboard]
Once eligible, the page shows:
### Your Referral Link [#your-referral-link]
A unique shareable link tied to your organization. Click the copy button to copy it to your clipboard and share it with others.
### Your Stats [#your-stats]
| Stat | Description |
| ------------------ | ----------------------------------------------------- |
| **Users Referred** | Total number of users who signed up through your link |
| **Total Earnings** | Total credit amount earned from referrals |
### How It Works [#how-it-works]
1. **Share Your Link** — Send your referral link to others
2. **They Sign Up** — They create an Vichar account using your link
3. **Earn Credits** — You earn 1% of their spending as credits
Credits are automatically added to your organization balance.
# Self-Service Refunds
URL: https://docs.vichar.io/learn/refunds
Recent purchases can be refunded from the billing history without contacting support. The refund goes back to the original payment method.
## Where to find it [#where-to-find-it]
Every paid top-up in the billing history has a refund button. When the top-up can be refunded the button is active; when it cannot, the button is disabled and its tooltip says why.
Open **Billing → Transactions** in the organization settings.
## What can be refunded [#what-can-be-refunded]
| Purchase | Window | Conditions |
| ----------------- | ------- | --------------------------------------------------------------- |
| **Credit top-up** | 14 days | Your most recent top-up, with less than 20% of its credits used |
Only the organization owner can request a refund, and each purchase can be refunded once.
Refunding a top-up removes its credits from your balance once the refund
completes. Refunds cannot be undone.
## Requesting a refund [#requesting-a-refund]
1. Open the billing history for the product and find the charge.
2. Click **Request refund**.
3. Pick a reason. A short explanation is required when you choose **Other**.
4. Confirm with **Request refund**.
## Why a refund is unavailable [#why-a-refund-is-unavailable]
| Message | Meaning |
| --------------------------------------------------- | ----------------------------------- |
| Refunds are available for 14 days after purchase | The refund window has closed |
| Only your most recent purchase can be self-refunded | A newer top-up exists |
| More than 20% of these credits have been used | Usage is above the refund threshold |
| Only the organization owner can request a refund | Ask the owner to request it |
| This purchase has already been refunded | The charge was refunded before |
For anything these rules do not cover, contact support.
# Structured outputs
URL: https://docs.vichar.io/learn/structured-outputs
Vichar exposes **two distinct JSON capabilities**, and they are kept
separate on every surface: the models API, the [models
directory](https://app.vichar.io/dashboard), the model detail pages, and the
gateway's request validation.
## Soft JSON output (`json_output`) [#soft-json-output-json_output]
A model with soft JSON output can be **nudged** into emitting JSON — typically
via `response_format: { "type": "json_object" }` or prompt guidance. There is
no server-side schema guarantee: the model may still produce off-schema or
malformed JSON, so the consumer's parser must be able to handle it.
## Strict JSON output schema (`structured_outputs`) [#strict-json-output-schema-structured_outputs]
A model with strict JSON output schema is served by an **upstream provider
that natively enforces schema-guided decoding** (for example OpenAI structured
outputs or vLLM guided decoding). You provide a JSON schema via
`response_format: { "type": "json_schema", "json_schema": { ... } }` and the
provider guarantees the output conforms to it.
The gateway only **declares** this capability — it never emulates schema
enforcement (no prompt-and-validate adapter). A model whose provider does not
natively support `json_schema` rejects such requests with `400 does not
support JSON schema output mode`. That rejection is **by design**, not a
missing feature.
## Field mapping [#field-mapping]
| Models API field | Catalogue flag | Meaning |
| -------------------- | ------------------ | --------------------------------------------- |
| `json_output` | `jsonOutput` | Soft: nudged JSON, no schema guarantee |
| `structured_outputs` | `jsonOutputSchema` | Strict: upstream provider enforces the schema |
The snake\_case names are what the OpenAI-compatible [`/v1/models`](https://docs.vichar.io/learn/models) endpoint exposes;
the camelCase names are used in the model catalogue and the directory UI.
The two tiers are **independent per provider mapping**: a provider may support
strict `json_schema` without soft `json_object` (for example Runware or
Perplexity declare `structured_outputs` without `json_output`), or soft
`json_object` without strict schema. The gateway routes each mode by its own
flag — `json_schema` requests only require `structured_outputs`, they are never
emulated or silently downgraded.
## How to verify a model [#how-to-verify-a-model]
Probe a model with a strict `json_schema` request:
```bash
curl -s https://api.vichar.io/v1/chat/completions \
-H "Authorization: Bearer $LLMGATEWAY_API_KEY" \
-d '{
"model": "gpt-4o-mini",
"messages": [{ "role": "user", "content": "Reply JSON: {\"ok\":true}" }],
"max_tokens": 16,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "probe",
"strict": true,
"schema": {
"type": "object",
"properties": { "ok": { "type": "boolean" } },
"required": ["ok"],
"additionalProperties": false
}
}
}
}'
```
* `200` → the upstream provider enforces the schema; the model should declare `structured_outputs: true`.
* `400 does not support JSON schema output mode` → soft-only; the model should expose `json_output` without `structured_outputs`.
The probe above can be run for any model listed in `GET /v1/models`; a
declared flag that does not match the observed status is a catalog bug and
should be reported against the repository.
# Team
URL: https://docs.vichar.io/learn/team
The Team page has two views: **Members** for organization membership and personal controls, and **Teams** for policy shared by groups of developers.
## Members [#members]
The Members view shows each person's role, assigned organization team, projects, and limits. Enterprise owners and admins can also compare cost, tokens, requests, and API-key counts for the selected period.
Each row has actions to:
* Open the member's details and usage
* Change their role and project access
* Set a personal API-key or spending budget
* Add personal IAM rules
* Remove them from the organization
Personal IAM rules apply to every key created by that member. A key can narrow those rules, but cannot expand access beyond them. See [Member-level IAM rules](https://docs.vichar.io/features/api-keys#member-level-iam-rules) for rule behavior.
### Roles [#roles]
| Role | Permissions |
| ----------------- | --------------------------------------------------------------------------------------------------- |
| **Owner** | Full access, including team management, billing, and organization settings |
| **Admin** | Can manage members, projects, and API keys, but cannot change billing controls or modify owners |
| **Project admin** | Manages settings, routing, keys, and usage in assigned projects; cannot administer the organization |
| **Developer** | Creates and manages their own API keys and views their own usage in assigned projects |
### Invite a member [#invite-a-member]
Click **Add Member**, enter an email address, and choose a role. Project admins and developers require Enterprise access and at least one project grant. The invitation remains under **Pending Invitations** until it is accepted or revoked.
Use **Project admin** for someone who needs to manage a project's settings and all its API keys and usage. Use **Manage access** on an existing member to change their role or replace their project grants. Organization settings, billing controls, provider keys, and membership remain restricted to organization administrators. See [Project access](https://docs.vichar.io/features/project-access) for the complete permission model.
Pending invitations reserve a seat. The Members card shows the current seat count and plan limit, and **Add Member** is disabled when the limit is reached.
## Organization teams [#organization-teams]
Open **Teams** to group developers under a shared project, IAM, and budget policy. One developer can belong to one organization team; other roles cannot be assigned. Promoting a Developer to Project admin clears their team assignment.
Creating teams and changing team policy requires the [**Enterprise
plan**](mailto:contact@vichar.io). If Enterprise access ends, existing policy
remains enforced and can still be reviewed. Developers can be unassigned, and
empty teams can be deleted.
Click **Create team**, give it a unique name, then click **Open** to configure it. The list shows the number of developers, project ceiling, and IAM rule count for every team.
## Configure team policy [#configure-team-policy]
| Control | Behavior |
| --------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| **Team identity** | Names the team and can mark it as the organization's **default team**. Names are unique within the organization. |
| **Project ceiling** | Effective access is the intersection of the team list and each developer's personal grants. With no projects selected, user keys return 403. |
| **Per-developer budget ceilings** | Caps each developer's active keys, lifetime spend, and recurring spend. Personal and API-key limits may be stricter. |
| **Developers** | Assigns, moves, or unassigns Developers. A move replaces the previous team's policy immediately. |
| **IAM policy** | Adds shared model, provider, pricing, or IP rules. Member and API-key rules run afterward and can only narrow access. |
Unassigning a developer restores their personal project, IAM, and budget settings. Unassign every developer before deleting a team.
### Default team [#default-team]
One team per organization can be marked as the default. Developers who join — by invite or direct add — are assigned to it automatically; explicit manual assignments always win. When enabling the default you can optionally assign every developer currently without a team in one step. Moving or removing the flag only changes future joins: current members keep their team.
# Transactions
URL: https://docs.vichar.io/learn/transactions
The Transactions page shows a complete history of all financial transactions in your organization.
## Transaction History [#transaction-history]
Each transaction entry includes:
| Field | Description |
| --------------- | ---------------------------------------- |
| **Date** | When the transaction occurred |
| **Type** | The transaction type (see below) |
| **Credits** | Number of credits added or deducted |
| **Total Paid** | The dollar amount charged |
| **Status** | Current state of the transaction |
| **Description** | Additional details about the transaction |
## Transaction Types [#transaction-types]
| Type | Description |
| ----------------- | ----------------------------------- |
| **Credit Top-up** | Manual or automatic credit purchase |
| **Credit Refund** | Credits deducted for a refund |
| **Credit Gift** | Credits granted by support |
| **Credit Usage** | Credits consumed by API usage |
## Status Badges [#status-badges]
* **Completed** — Transaction processed successfully
* **Pending** — Transaction is being processed
* **Failed** — Transaction could not be completed
## Refunds [#refunds]
Recent top-ups can be refunded from this page with **Request refund**. See [Self-Service Refunds](https://docs.vichar.io/learn/refunds) for the window and conditions.
# Usage & Metrics
URL: https://docs.vichar.io/learn/usage-metrics
The Usage & Metrics page provides comprehensive analytics through five tabs, giving you deep insight into your LLM API usage patterns.
## Filters [#filters]
* **API Key** — Filter metrics by a specific API key or view all
* **Date range** — Select the time period (defaults to last 7 days)
## Tabs [#tabs]
### Requests [#requests]
A time-series chart showing request volume over the selected period. Use this to identify traffic patterns, peak usage times, and growth trends.
### Models [#models]
A table showing your top-used models ranked by request count. For each model you can see:
* Total requests
* Token consumption
* Associated costs
This helps you understand which models drive the most usage and cost.
### Errors [#errors]
A chart showing error rates over time. Track:
* Error frequency and trends
* Spikes that may indicate provider issues
* Overall reliability of your API calls
### Cache [#cache]
A chart showing your cache hit rate over time. Monitor:
* How effectively caching is reducing redundant requests
* Cache hit vs. miss ratios
* The cost savings from cached responses
### Costs [#costs]
A cost breakdown chart showing spending patterns. Analyze:
* Cost trends over time
* Cost distribution by provider or model
* Opportunities to reduce spending
# Migrate from LiteLLM
URL: https://docs.vichar.io/migrations/litellm
Running your own LiteLLM proxy works—until it doesn't. Scaling, monitoring, and keeping it running becomes another job. Vichar gives you the same unified API with built-in analytics, caching, and a dashboard—without the infrastructure overhead.
## Quick Migration [#quick-migration]
Both services use OpenAI-compatible endpoints, so migration is a two-line change:
```diff
- const baseURL = "http://localhost:4000/v1"; // LiteLLM proxy
+ const baseURL = "https://api.vichar.io/v1";
- const apiKey = process.env.LITELLM_API_KEY;
+ const apiKey = process.env.LLM_GATEWAY_API_KEY;
```
## Why Teams Switch to Vichar [#why-teams-switch-to-vichar]
| What You Get | LiteLLM (Self-Hosted) | Vichar |
| ------------------------ | --------------------- | -------------------- |
| OpenAI-compatible API | Yes | Yes |
| Infrastructure to manage | Yes (you run it) | No (we run it) |
| Managed cloud option | No | Yes |
| Analytics dashboard | Basic | Per-request detail |
| Response caching | Manual setup | Built-in, automatic |
| Cost tracking | Via callbacks | Native, real-time |
| Provider key management | Config file | Web UI with rotation |
| Uptime & scaling | You handle it | 99.9% SLA (managed) |
Self-hosting also means you own the patch cycle. On March 24, 2026 two malicious `litellm` releases (1.82.7 and 1.82.8) reached PyPI and were pulled after about 40 minutes; they harvested credentials from unpinned pip installs, while pinned Docker deployments were unaffected. Pin versions wherever you self-host — Vichar included — or use the managed gateway and skip the upkeep.
Still want to self-host? Vichar is open source under AGPLv3 — see [Self-Hosting](https://docs.vichar.io/self-host/docker-compose)—same features, your infrastructure.
For a detailed breakdown, see [Getting Started](https://docs.vichar.io/quick-start).
## Migration Steps [#migration-steps]
### Get Your Vichar API Key [#get-your-vichar-api-key]
Sign up at [app.vichar.io/signup](https://app.vichar.io/signup) and create an API key from your dashboard.
### Map Your Models [#map-your-models]
Vichar supports two model ID formats:
**Canonical Model IDs** (without provider prefix) - Uses smart routing to automatically select the best provider based on uptime, throughput, price, and latency:
```
gpt-6-astra
claude-sonnet-5
gemini-3.1-pro-preview
```
**Provider-Prefixed Model IDs** - Routes to a specific provider with automatic failover if uptime drops below 90%:
```
openai/gpt-6-astra
anthropic/claude-sonnet-5
google-ai-studio/gemini-3.1-pro-preview
```
This means many LiteLLM model names work directly with Vichar:
| LiteLLM Model | Vichar Model |
| --------------------------------- | ----------------------------------------------------------------- |
| gpt-6-astra | gpt-6-astra or openai/gpt-6-astra |
| anthropic/claude-sonnet-5 | claude-sonnet-5 or anthropic/claude-sonnet-5 |
| gemini/gemini-3.1-pro-preview | gemini-3.1-pro-preview or google-ai-studio/gemini-3.1-pro-preview |
| bedrock/anthropic.claude-sonnet-5 | claude-sonnet-5 or aws-bedrock/claude-sonnet-5 |
For more details on routing behavior, see the [routing documentation](https://docs.vichar.io/features/routing).
### Update Your Code [#update-your-code]
#### Python with OpenAI SDK [#python-with-openai-sdk]
```python
from openai import OpenAI
# Before (LiteLLM proxy)
client = OpenAI(
base_url="http://localhost:4000/v1",
api_key=os.environ["LITELLM_API_KEY"]
)
response = client.chat.completions.create(
model="gpt-6-astra",
messages=[{"role": "user", "content": "Hello!"}]
)
# After (Vichar) - model name can stay the same!
client = OpenAI(
base_url="https://api.vichar.io/v1",
api_key=os.environ["LLM_GATEWAY_API_KEY"]
)
response = client.chat.completions.create(
model="gpt-6-astra", # or "openai/gpt-6-astra" to target a specific provider
messages=[{"role": "user", "content": "Hello!"}]
)
```
#### Python with LiteLLM Library [#python-with-litellm-library]
If you're using the LiteLLM library directly, you can point it to Vichar:
```python
import litellm
# Before (direct LiteLLM)
response = litellm.completion(
model="gpt-6-astra",
messages=[{"role": "user", "content": "Hello!"}]
)
# After (via Vichar) - same model name works
response = litellm.completion(
model="gpt-6-astra", # or "openai/gpt-6-astra" to target a specific provider
messages=[{"role": "user", "content": "Hello!"}],
api_base="https://api.vichar.io/v1",
api_key=os.environ["LLM_GATEWAY_API_KEY"]
)
```
#### TypeScript/JavaScript [#typescriptjavascript]
```typescript
import OpenAI from "openai";
// Before (LiteLLM proxy)
const client = new OpenAI({
baseURL: "http://localhost:4000/v1",
apiKey: process.env.LITELLM_API_KEY,
});
// After (Vichar) - same model name works
const client = new OpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const completion = await client.chat.completions.create({
model: "gpt-6-astra", // or "openai/gpt-6-astra" to target a specific provider
messages: [{ role: "user", content: "Hello!" }],
});
```
#### cURL [#curl]
```bash
# Before (LiteLLM proxy)
curl http://localhost:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"messages": [{"role": "user", "content": "Hello!"}]
}'
# After (Vichar) - same model name works
curl https://api.vichar.io/v1/chat/completions \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"messages": [{"role": "user", "content": "Hello!"}]
}'
# Use "openai/gpt-6-astra" to target a specific provider
```
### Migrate Configuration [#migrate-configuration]
#### LiteLLM Config (Before) [#litellm-config-before]
```yaml
# litellm_config.yaml
model_list:
- model_name: gpt-6-astra
litellm_params:
model: openai/gpt-6-astra
api_key: sk-...
- model_name: claude-sonnet-5
litellm_params:
model: anthropic/claude-sonnet-5
api_key: sk-ant-...
```
#### Vichar (After) [#vichar-after]
With Vichar, you don't need a config file. Provider keys are managed in the web dashboard, or you can use the default Vichar keys.
## Streaming Support [#streaming-support]
Vichar supports streaming identically to LiteLLM:
```python
from openai import OpenAI
client = OpenAI(
base_url="https://api.vichar.io/v1",
api_key=os.environ["LLM_GATEWAY_API_KEY"]
)
stream = client.chat.completions.create(
model="openai/gpt-6-astra",
messages=[{"role": "user", "content": "Write a story"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
```
## Function/Tool Calling [#functiontool-calling]
Vichar supports function calling:
```python
from openai import OpenAI
client = OpenAI(
base_url="https://api.vichar.io/v1",
api_key=os.environ["LLM_GATEWAY_API_KEY"]
)
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
}
}]
response = client.chat.completions.create(
model="openai/gpt-6-astra",
messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
tools=tools
)
```
## Removing LiteLLM Infrastructure [#removing-litellm-infrastructure]
After verifying Vichar works for your use case, you can decommission your LiteLLM proxy:
1. Update all clients to use Vichar endpoints
2. Monitor the Vichar dashboard for successful requests
3. Shut down your LiteLLM proxy server
4. Remove LiteLLM configuration files
## What Changes After Migration [#what-changes-after-migration]
* **No servers to babysit** — We handle scaling, uptime, and updates
* **Real-time cost visibility** — See what every request costs, broken down by model
* **Automatic caching** — Repeated requests hit cache, reducing your spend
* **Web-based management** — No more editing YAML files for config changes
* **New models immediately** — Access new releases within 48 hours, no deployment needed
## Self-Hosting Vichar [#self-hosting-vichar]
If you prefer self-hosting like LiteLLM, Vichar is available under AGPLv3:
```bash
git clone https://github.com/vicharai/api
cd llmgateway
pnpm install
pnpm run setup
pnpm dev
```
This gives you the same benefits as LiteLLM's self-hosted proxy with Vichar's analytics and caching features.
## Full Comparison [#full-comparison]
Want to see a detailed breakdown of all features? See [Quick Start](https://docs.vichar.io/quick-start).
# Migrate from OpenRouter
URL: https://docs.vichar.io/migrations/openrouter
Vichar works just like OpenRouter—same OpenAI-compatible API, same `provider/model` naming—with built-in analytics and the option to self-host. Migration takes two lines of code.
Stripe announced its acquisition of OpenRouter on August 19, 2026. OpenRouter
says nothing changes for customers, so there is no fire to put out. This guide
is for teams that want an open-source gateway they can run themselves, or that
want to stop paying 5% on bring-your-own-key traffic above $25,000 a month.
## Quick Migration [#quick-migration]
Change your base URL and API key:
```diff
- const baseURL = "https://openrouter.ai/api/v1";
- const apiKey = process.env.OPENROUTER_API_KEY;
+ const baseURL = "https://api.vichar.io/v1";
+ const apiKey = process.env.LLM_GATEWAY_API_KEY;
```
## Migration Steps [#migration-steps]
### Get Your Vichar API Key [#get-your-vichar-api-key]
Sign up at [app.vichar.io/signup](https://app.vichar.io/signup) and create an API key from your dashboard.
### Update Environment Variables [#update-environment-variables]
```bash
# Remove OpenRouter credentials
# OPENROUTER_API_KEY=sk-or-...
# Add Vichar credentials
LLM_GATEWAY_API_KEY=vichar_your_key_here
```
### Update Your Code [#update-your-code]
#### Using fetch/axios [#using-fetchaxios]
The OpenRouter-only `HTTP-Referer` and `X-Title` headers can go; nothing on Vichar reads them.
```typescript
// Before (OpenRouter)
const response = await fetch("https://openrouter.ai/api/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}`,
"HTTP-Referer": "https://example.com",
"X-Title": "My App",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "openai/gpt-6-astra",
messages: [{ role: "user", content: "Hello!" }],
}),
});
// After (Vichar)
const response = await fetch("https://api.vichar.io/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LLM_GATEWAY_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "openai/gpt-6-astra",
messages: [{ role: "user", content: "Hello!" }],
}),
});
```
#### Using OpenAI SDK [#using-openai-sdk]
```typescript
import OpenAI from "openai";
// Before (OpenRouter)
const client = new OpenAI({
baseURL: "https://openrouter.ai/api/v1",
apiKey: process.env.OPENROUTER_API_KEY,
});
// After (Vichar)
const client = new OpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
// Usage remains the same
const completion = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5",
messages: [{ role: "user", content: "Hello!" }],
});
```
#### Using Vercel AI SDK [#using-vercel-ai-sdk]
Both OpenRouter and Vichar have native AI SDK providers, making migration straightforward:
```typescript
import { generateText } from "ai";
// Before (OpenRouter AI SDK Provider)
import { createOpenRouter } from "@openrouter/ai-sdk-provider";
const openrouter = createOpenRouter({
apiKey: process.env.OPENROUTER_API_KEY,
});
const { text } = await generateText({
model: openrouter("openai/gpt-6-astra"),
prompt: "Hello!",
});
// After (Vichar AI SDK Provider)
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { text } = await generateText({
model: llmgateway("openai/gpt-6-astra"),
prompt: "Hello!",
});
```
## Model Name Mapping [#model-name-mapping]
Most OpenRouter IDs work unchanged. A bare ID (no prefix) turns on smart routing across every provider that serves the model; a `provider/model` ID pins one provider. Anthropic versions use dashes instead of OpenRouter's dots.
| OpenRouter Model | Vichar Model |
| ----------------------------- | ----------------------------------------------------------------- |
| openai/gpt-6-astra | gpt-6-astra or openai/gpt-6-astra |
| anthropic/claude-sonnet-5 | claude-sonnet-5 or anthropic/claude-sonnet-5 |
| anthropic/claude-opus-4.8 | claude-opus-4-8 or anthropic/claude-opus-4-8 |
| google/gemini-3.1-pro-preview | gemini-3.1-pro-preview or google-ai-studio/gemini-3.1-pro-preview |
Check the [models page](https://app.vichar.io/dashboard) for the full list of available models.
## Provider Routing [#provider-routing]
OpenRouter selects the upstream through a `provider` object in the request body. Vichar puts that choice in the model ID and a header:
| OpenRouter | Vichar |
| ------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| No `provider` object | Bare model ID — routes on live uptime, throughput, price, and latency |
| `provider.order: ["Anthropic"]` | `anthropic/claude-sonnet-5` — pinned; falls back to the best alternative only if its uptime drops below 90% |
| `provider.order` with several entries | Provider routing preferences applied as an ordered fallback list |
| `provider.allow_fallbacks: false` | Add the `x-no-fallback: true` header to fail instead of retrying elsewhere |
See the [routing](https://docs.vichar.io/features/routing) documentation for the details.
## Streaming Support [#streaming-support]
Vichar supports streaming responses identically to OpenRouter:
```typescript
const stream = await client.chat.completions.create({
model: "anthropic/claude-sonnet-5",
messages: [{ role: "user", content: "Write a story" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}
```
## Full Comparison [#full-comparison]
Want to see a detailed breakdown of all features? See [Quick Start](https://docs.vichar.io/quick-start).
# Migrate from Vercel AI Gateway
URL: https://docs.vichar.io/migrations/vercel-ai-gateway
## Quick Migration [#quick-migration]
Swap your provider imports—your AI SDK code stays the same:
```diff
- import { openai } from "@ai-sdk/openai";
- import { anthropic } from "@ai-sdk/anthropic";
+ import { generateText } from "ai";
+ import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
+ const llmgateway = createLLMGateway({
+ apiKey: process.env.LLM_GATEWAY_API_KEY
+ });
const { text } = await generateText({
- model: openai("gpt-6-astra"),
+ model: llmgateway("gpt-6-astra"),
prompt: "Hello!"
});
```
The key difference: one provider, one API key, all models—with caching and analytics built in.
## Zero-diff alternative: repoint the base URL [#zero-diff-alternative-repoint-the-base-url]
If your app passes bare model strings (`model: "anthropic/claude-sonnet-5"`), it resolves them through `@ai-sdk/gateway` — the AI SDK's default provider. Vichar implements that protocol, so you can keep every line of application code and repoint the provider instead:
```ts
import { createGateway } from "@ai-sdk/gateway";
globalThis.AI_SDK_DEFAULT_PROVIDER = createGateway({
baseURL: "https://api.vichar.io/v4/ai",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
```
No import changes, no model-string changes, and provider-native web search keeps returning `source-url` parts. See [AI SDK Gateway protocol](https://docs.vichar.io/developers/ai-sdk-gateway-protocol).
Prefer the explicit `@llmgateway/ai-sdk-provider` migration below when you want the gateway's own model IDs and options surfaced as first-class provider settings.
## Migration Steps [#migration-steps]
### Get Your Vichar API Key [#get-your-vichar-api-key]
Sign up at [app.vichar.io/signup](https://app.vichar.io/signup) and create an API key from your dashboard.
### Install the Vichar AI SDK Provider [#install-the-vichar-ai-sdk-provider]
Install the native Vichar provider for the Vercel AI SDK:
```bash
pnpm add @llmgateway/ai-sdk-provider
```
This package provides full compatibility with the Vercel AI SDK and supports all Vichar features.
### Update Your Code [#update-your-code]
#### Basic Text Generation [#basic-text-generation]
```typescript
// Before (Vercel AI Gateway with native providers)
import { openai } from "@ai-sdk/openai";
import { anthropic } from "@ai-sdk/anthropic";
import { generateText } from "ai";
const { text: openaiText } = await generateText({
model: openai("gpt-6-astra"),
prompt: "Hello!",
});
const { text: claudeText } = await generateText({
model: anthropic("claude-sonnet-5"),
prompt: "Hello!",
});
// After (Vichar - single provider for all models)
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { generateText } from "ai";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { text: openaiText } = await generateText({
model: llmgateway("openai/gpt-6-astra"),
prompt: "Hello!",
});
const { text: claudeText } = await generateText({
model: llmgateway("anthropic/claude-sonnet-5"),
prompt: "Hello!",
});
```
#### Streaming Responses [#streaming-responses]
```typescript
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { streamText } from "ai";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { textStream } = await streamText({
model: llmgateway("anthropic/claude-sonnet-5"),
prompt: "Write a poem about coding",
});
for await (const text of textStream) {
process.stdout.write(text);
}
```
#### Using in Next.js API Routes [#using-in-nextjs-api-routes]
```typescript
// app/api/chat/route.ts
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { streamText } from "ai";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
export async function POST(req: Request) {
const { messages } = await req.json();
const result = await streamText({
model: llmgateway("openai/gpt-6-astra"),
messages,
});
return result.toUIMessageStreamResponse();
}
```
#### Alternative: Using OpenAI SDK Adapter [#alternative-using-openai-sdk-adapter]
If you prefer not to install a new package, you can use `@ai-sdk/openai` with a custom base URL:
```typescript
import { createOpenAI } from "@ai-sdk/openai";
import { generateText } from "ai";
const llmgateway = createOpenAI({
baseURL: "https://api.vichar.io/v1",
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { text } = await generateText({
model: llmgateway("openai/gpt-6-astra"),
prompt: "Hello!",
});
```
### Update Environment Variables [#update-environment-variables]
```bash
# Remove individual provider keys (optional - can keep as backup)
# OPENAI_API_KEY=sk-...
# ANTHROPIC_API_KEY=sk-ant-...
# Add Vichar key
export LLM_GATEWAY_API_KEY=vichar_your_key_here
```
## Model Name Format [#model-name-format]
Vichar supports two model ID formats:
**Canonical Model IDs** (without provider prefix) - Uses smart routing to automatically select the best provider based on uptime, throughput, price, and latency:
```
gpt-6-astra
claude-sonnet-5
gemini-3.1-pro-preview
```
**Provider-Prefixed Model IDs** - Routes to a specific provider with automatic failover if uptime drops below 90%:
```
openai/gpt-6-astra
anthropic/claude-sonnet-5
google-ai-studio/gemini-3.1-pro-preview
```
For more details on routing behavior, see the [routing documentation](https://docs.vichar.io/features/routing).
### Model Mapping Examples [#model-mapping-examples]
| Vercel AI SDK | Vichar |
| ---------------------------------- | ------------------------------------------------------------------------------------------------- |
| `openai("gpt-6-astra")` | `llmgateway("gpt-6-astra")` or `llmgateway("openai/gpt-6-astra")` |
| `anthropic("claude-sonnet-5")` | `llmgateway("claude-sonnet-5")` or `llmgateway("anthropic/claude-sonnet-5")` |
| `google("gemini-3.1-pro-preview")` | `llmgateway("gemini-3.1-pro-preview")` or `llmgateway("google-ai-studio/gemini-3.1-pro-preview")` |
Check the [models page](https://app.vichar.io/dashboard) for the full list of available models.
## Tool Calling [#tool-calling]
Vichar supports tool calling through the AI SDK:
```typescript
import { createLLMGateway } from "@llmgateway/ai-sdk-provider";
import { generateText, tool } from "ai";
import { z } from "zod";
const llmgateway = createLLMGateway({
apiKey: process.env.LLM_GATEWAY_API_KEY,
});
const { text, toolResults } = await generateText({
model: llmgateway("openai/gpt-6-astra"),
tools: {
weather: tool({
description: "Get the weather for a location",
inputSchema: z.object({
location: z.string(),
}),
execute: async ({ location }) => {
return { temperature: 72, condition: "sunny" };
},
}),
},
prompt: "What's the weather in San Francisco?",
});
```
## Self-Hosting Vichar [#self-hosting-vichar]
If you prefer self-hosting, Vichar is available under AGPLv3:
```bash
git clone https://github.com/vicharai/api
cd llmgateway
pnpm install
pnpm run setup
pnpm dev
```
This gives you the same managed experience with full control over your infrastructure.
# Vichar API Versioning and Deprecation Policy
URL: https://docs.vichar.io/resources/api-versioning
Vichar's public REST API uses URL path versions. The current OpenAI-compatible base URL is `https://api.vichar.io/v1`. The [OpenAPI specification](https://api.vichar.io/openapi.json) describes the supported operations and typed errors.
## Compatibility [#compatibility]
Breaking changes to the REST contract use a new path version. Additive fields, new endpoints, and new model or provider mappings can appear within an existing version. Clients should tolerate unknown response fields.
The AI SDK gateway protocol has its own versioned paths and negotiation. MCP negotiates its protocol version during initialization; neither uses the REST path version as its protocol version.
## Deprecation and sunset [#deprecation-and-sunset]
API deprecation notices are published in the [docs](https://docs.vichar.io), with the affected surface, replacement, migration steps, and any scheduled retirement date. Deprecation means a surface is discouraged; sunset means it will stop serving requests. No retirement date is currently scheduled for `/v1`.
When an endpoint is deprecated, its notice defines the transition period. There is no blanket minimum notice period. Model availability can change on an upstream provider's schedule, independently of API versioning.
Deprecated endpoints use the [Deprecation header](https://www.rfc-editor.org/rfc/rfc9745.html) with an HTTP Structured Fields date and a `Link` with `rel="deprecation"` pointing to the notice. A scheduled retirement additionally uses the [Sunset header](https://www.rfc-editor.org/rfc/rfc8594.html) with an HTTP date. These headers are omitted on endpoints without an announced deprecation or retirement.
Deprecation notices are published in the docs.
Clients should monitor the changelog, migrate before a published sunset, and treat a missing `Sunset` header as an unspecified date.
## Model and provider lifecycle [#model-and-provider-lifecycle]
Model and provider availability is separate from the REST contract. Consult the live [models](https://app.vichar.io/dashboard) and [providers](https://app.vichar.io/dashboard) directories before choosing an integration target. Deprecated entries remain available for historical lookups; deactivated mappings cannot serve new requests.
# Vichar API Errors
URL: https://docs.vichar.io/resources/error-handling
On the OpenAI-compatible endpoints, Vichar returns errors in the same format as the OpenAI API, so existing OpenAI SDKs and tooling can parse gateway errors without changes. This applies to errors forwarded from upstream providers as well as errors raised by the gateway itself (authentication failures, usage limits, validation problems, timeouts, and so on). The Anthropic-compatible Messages endpoint (`/v1/messages`) instead returns Anthropic-native errors — see [Anthropic Endpoint](#anthropic-endpoint) below.
Upstream rate limits are treated as provider errors and are eligible for retries and fallback. Rate-limit headers returned to clients describe only limits enforced by Vichar; upstream `Retry-After` and rate-limit headers are not forwarded.
## Error Format [#error-format]
Errors on the OpenAI-compatible endpoints (`/v1/chat/completions`, `/v1/embeddings`, `/v1/images`, `/v1/models`, `/v1/moderations`, `/v1/rerank`, `/v1/responses`, `/v1/videos`) use the standard OpenAI error envelope:
```json
{
"error": {
"message": "Unauthorized: Vichar API key reached its usage limit.",
"type": "invalid_request_error",
"param": null,
"code": "invalid_api_key"
}
}
```
| Field | Description |
| --------------- | ----------------------------------------------------------------------------------- |
| `error.message` | Human-readable description of what went wrong. |
| `error.type` | High-level error category (see the table below). |
| `error.param` | The request parameter that caused the error, or `null` when not parameter-specific. |
| `error.code` | A more specific machine-readable code, or `null` when no specific code applies. |
For normal JSON responses, the HTTP status code matches the error and is the authoritative signal — read it from the response status line rather than the body. The exception is an error arriving mid-stream, where the status is already `200` and the error must be read from the SSE event payload (see [Streaming Errors](#streaming-errors)).
## Status Codes [#status-codes]
For **gateway-raised** errors (authentication failures, usage limits, validation problems, rate limits, timeouts), the gateway maps HTTP status codes to OpenAI error types and codes as follows:
| Status | `type` | `code` |
| ------ | ----------------------- | ------------------------ |
| 400 | `invalid_request_error` | *(varies / `null`)* |
| 401 | `invalid_request_error` | `invalid_api_key` |
| 402 | `invalid_request_error` | `billing_error` |
| 403 | `invalid_request_error` | `permission_denied` |
| 404 | `invalid_request_error` | `not_found` |
| 408 | `timeout_error` | `timeout` |
| 410 | `invalid_request_error` | *(varies / `null`)* |
| 413 | `invalid_request_error` | `request_too_large` |
| 415 | `invalid_request_error` | `unsupported_media_type` |
| 429 | `rate_limit_error` | `rate_limit_exceeded` |
| 499 | `invalid_request_error` | `request_cancelled` |
| 504 | `timeout_error` | `timeout` |
| 529 | `overloaded` | `overloaded` |
| 5xx | `api_error` | *(`null`)* |
Validation errors raised before a request reaches a provider often include a
more specific `code` and a `param` pointing at the offending field — for
example `invalid_json`, `model_not_found`, or
`unsupported_parameter_combination`.
## Upstream Provider Errors [#upstream-provider-errors]
Errors from the upstream provider follow different rules on the chat completions path:
* **Upstream 4xx client errors** (a genuinely invalid request — any 4xx other than 401/402/403/404/405/429, or a 400 whose body signals a provider-side problem such as bad credentials, an exhausted provider account, or an unknown model — those are treated as provider-side failures) are passed through with their **original status code**. An already OpenAI-shaped error body is forwarded unchanged; bare provider shapes are wrapped in the OpenAI envelope with `type` and `code` set to `client_error` (or the provider's own error type), and only these wrapped bodies carry extra diagnostic fields inside `error`: `requestedProvider`, `usedProvider`, `requestedModel`, `usedInternalModel`, and `responseText`.
* **Upstream provider/gateway-side failures** (5xx, upstream 429/404, provider credential or funding problems) first go through [automatic retry and fallback](https://docs.vichar.io/features/routing#automatic-retry--fallback). If no provider succeeds, the gateway returns **HTTP 500** with `type`/`code` of `upstream_error` or `gateway_error` and the same diagnostic fields — the upstream status code is not passed through (it is recorded in the request log's error details). Some exhaustion paths instead return **HTTP 502** with `type` `upstream_error` and `code` `all_providers_failed`, without the diagnostic fields.
* **Connection failures and upstream timeouts** return **HTTP 502** (`upstream_error`/`fetch_failed`) or **HTTP 504** (`upstream_timeout`/`timeout`).
## Blocked Accounts [#blocked-accounts]
When an administrator blocks an account, gateway requests return HTTP `410`.
If a reason was provided, it appears in `error.message`. The same reason is shown
when affected members try to sign in (HTTP `403`) and in authenticated API errors.
Older blocks and blocks without a reason retain the generic error message.
Blocking also signs out all members and cancels the organization's subscriptions.
Re-enabling the organization does not reactivate its members or restore subscriptions.
## Streaming Errors [#streaming-errors]
For streaming requests (`"stream": true`), an error that occurs **after** the stream has started is delivered as an SSE `error` event whose payload uses the same `{ "error": { ... } }` envelope. Because the HTTP status is already `200` at that point, read the error from the event payload; its `type`/`code` values are gateway-specific (for example `upstream_timeout`/`timeout`, `upstream_error`/`all_providers_failed`, or `gateway_error`) rather than the status-code table above, and `param` may be absent. Errors that occur **before** streaming begins (such as authentication failures) are returned as a normal JSON error response with the appropriate status code.
## Gateway Content Filter [#gateway-content-filter]
Requests routed to some providers may be screened by Vichar's own content filter, which runs the prompt (and any image inputs) through a moderation model before the request reaches the provider. The filter runs on a sample of requests and its strictness follows your organization's [trust tier](https://docs.vichar.io/resources/rate-limits#trust-tiers-account-age-or-spend): tiers 0–2 use strict thresholds, tiers 3–4 lenient ones. Vichar staff can pin an organization's content filter tier independently of its trust tier.
By default the filter only records its findings. When blocking is active, a request over its tier's thresholds is rejected before dispatch and the response makes clear that the gateway filter, not the provider, stopped it:
* **Chat Completions** (and the Responses endpoint): HTTP `200` with `finish_reason: "content_filter"`. The assistant message content explains that the request was blocked by Vichar's content filter and how to contact support. Streaming responses send one chunk carrying that content and finish reason, then `[DONE]`.
* **Messages** (`/v1/messages`): `stop_reason: "refusal"` with the same explanation as the text content.
* **Images**: HTTP `200` with an empty `data` array.
* **Videos**: HTTP `403` with the explanation in `error.message`; no job is created.
A moderation outage never fails a request: if the moderation model is unavailable, the request goes through as normal.
An occasional block is not a problem. From October 15, 2026, a high rate or volume of content filter violations can get the account rate-limited, restricted, suspended, or terminated at our discretion — see Section 6 of the [Terms of Use](https://app.vichar.io/legal/terms).
[**Enterprise**](mailto:contact@vichar.io) organizations are exempt from
blocking by default; they are only ever blocked if Vichar staff enable
enterprise enforcement. If a legitimate request is blocked, email
[contact@vichar.io](mailto:contact@vichar.io) so we can review your use case
against the Terms of Use and help you get unblocked.
## Anthropic Endpoint [#anthropic-endpoint]
The Anthropic-compatible Messages endpoint (`/v1/messages`) returns errors in Anthropic's native format instead, so the Anthropic SDK can parse them:
```json
{
"type": "error",
"error": {
"type": "authentication_error",
"message": "Unauthorized: invalid API key."
}
}
```
## Related [#related]
* [Rate Limits](https://docs.vichar.io/resources/rate-limits) — details on `429` responses and rate limit headers.
# Rate Limits
URL: https://docs.vichar.io/resources/rate-limits
Vichar applies rate limits to ensure fair usage and protect platform stability. Limits are evaluated in a few independent layers:
* **Per-organization endpoint limits** — a requests-per-minute cap on every API endpoint, scoped to your organization.
* **Free model limits** — additional limits specifically for zero-cost models.
* **Provider/model caps** — limits configured and enforced by Vichar for a provider/model.
## Per-Organization Endpoint Limits [#per-organization-endpoint-limits]
Every API endpoint is rate limited per organization using a rolling 60-second window. The limit is independent for each endpoint, so traffic to `/v1/chat/completions` does not consume the budget for `/v1/embeddings`.
The default limits (requests per minute, per organization) are:
| Endpoint | Path | Requests / min |
| -------------------- | -------------------------- | -------------- |
| Chat completions | `/v1/chat/completions` | 600 |
| Messages (Anthropic) | `/v1/messages` | 600 |
| Responses | `/v1/responses` | 600 |
| Embeddings | `/v1/embeddings` | 1200 |
| Moderations | `/v1/moderations` | 1200 |
| Rerank | `/v1/rerank` | 1200 |
| System One | `/v1/systemone` | 600 |
| Models | `/v1/models` | 1200 |
| OCR | `/v1/ocr` | 300 |
| Images | `/v1/images` | 300 |
| Speech | `/v1/audio/speech` | 300 |
| Transcriptions | `/v1/audio/transcriptions` | 300 |
| Videos | `/v1/videos` | 120 |
| Realtime (mint) | `/v1/realtime` | 120 |
| Key | `/v1/key` | 1200 |
| Credits | `/v1/credits` | 300 |
| AI SDK protocol | `/v*/ai` | 600 |
**[Enterprise](mailto:contact@vichar.io) organizations are exempt** from these
per-organization endpoint limits. [Contact us](mailto:contact@vichar.io) about
enterprise plans.
### Trust Tiers (account age or spend) [#trust-tiers-account-age-or-spend]
For regular (pay-as-you-go) organizations, limits scale with a **trust tier**. An organization qualifies for a tier when its account is old enough, **or** when its lifetime usage spend is high enough **and** the account meets the tier's minimum age. The tier raises the per-endpoint RPM limits, the [concurrent-request ceiling](#concurrent-request-limits), **and** the daily/monthly USD spend caps below.
| Tier | Qualifies (age, or spend + min age) | RPM multiplier | Concurrent | Daily cap | Monthly cap |
| ---- | ----------------------------------- | -------------- | ---------- | --------- | ----------- |
| 0 | new / $0 | 1× | 100 | $25 | $250 |
| 1 | 7 days **or** $10 (account ≥ 1 day) | 2× | 200 | $100 | $1,000 |
| 2 | 30 days **or** $100 (≥ 3 days) | 4× | 400 | $500 | $5,000 |
| 3 | 60 days **or** $1,000 (≥ 7 days) | 10× | 1,000 | $5,000 | $50,000 |
| 4 | 90 days **or** $5,000 (≥ 14 days) | 20× | 2,000 | $15,000 | $200,000 |
Spend alone never promotes a brand-new account: each spend-qualified tier also requires the minimum account age shown, so the fastest possible path to Tier 4 is 14 days — no amount of day-one usage unlocks higher limits.
The trust tier also selects how strict the [gateway content filter](https://docs.vichar.io/resources/error-handling#gateway-content-filter) is for your organization: tiers 0–2 are screened with strict thresholds, tiers 3–4 with lenient ones.
For example, an org past 30 days old (or with $100+ of usage) is Tier 2: chat completions rises from 600 to 2,400 RPM, with a $500/day and $5,000/month spend ceiling.
Qualifying spend counts **usage billed to your credit balance only** — usage served through your own provider keys (BYOK) does not count — and is **net of refunds**: every refunded payment is deducted, so refunded or clawed-back money never raises limits. Refunded top-ups still count against the top-up allowance below — refunding does not free up top-up headroom.
### Daily & Monthly Spend Caps [#daily--monthly-spend-caps]
Regular organizations also have hard **USD spend ceilings** — a daily and a monthly cap set by the trust tier above — so a brand-new account has a tight dollar velocity limit that rises as it ages or spends. Only real paid usage counts; **free models are exempt**, as is **enterprise** (no caps). When a cap is reached, requests return `429` until the counter resets (UTC midnight for daily, first of the month for monthly).
### Top-Up Limits [#top-up-limits]
Credit top-ups are also velocity-limited by trust tier: each organization can add at most a tier-scaled gross USD amount to its balance per **rolling 24-hour window**. This applies before any charge is made — a top-up attempt over the allowance is rejected with `429` and no card is charged.
| Tier | Top-up allowance (rolling 24h) |
| ---- | ------------------------------ |
| 0 | $100 |
| 1 | $500 |
| 2 | $2,500 |
| 3 | $10,000 |
| 4 | $20,000 |
The limit covers dashboard top-ups (card and hosted checkout) and auto top-up. **[Enterprise](mailto:contact@vichar.io) organizations are exempt** — [contact us](mailto:contact@vichar.io) if you need a higher allowance. Your current allowance and usage are shown on the Settings → Limits page. Hosted checkout links expire after 30 minutes.
### Enterprise [#enterprise]
Organizations on the **[Enterprise plan](mailto:contact@vichar.io) have no per-organization requests-per-minute, spend, or top-up limits**. Your request rate is limited only by your credit balance and any upstream provider limits. The only gateway limit that still applies is a greatly elevated [concurrent-request ceiling](#concurrent-request-limits).
Need unlimited gateway throughput? [Contact us](mailto:contact@vichar.io)
about an enterprise plan.
## Concurrent Request Limits [#concurrent-request-limits]
Separately from the per-minute request limits above, each organization has one fleet-wide budget of **concurrent in-flight requests** across all inference endpoints (chat completions, messages, responses, embeddings, moderations, rerank, System One, OCR, images, speech, transcriptions, videos, and the AI SDK surface). A slot is held for a request's full lifetime — including the entire duration of a streamed response — and freed when the response finishes or the connection closes.
This bounds what a per-minute limit cannot: long-running requests. Six hundred requests per minute that each stream for two minutes hold 1,200 connections open; the concurrency budget is what keeps that pile-up from exhausting shared gateway capacity.
For regular (pay-as-you-go) organizations the ceiling scales with the same [trust tier](#trust-tiers-account-age-or-spend) that raises the per-minute limits.
| Plan | Concurrent requests |
| ----------------------- | ------------------- |
| Regular (PAYG) — Tier 0 | 100 |
| Regular (PAYG) — Tier 1 | 200 |
| Regular (PAYG) — Tier 2 | 400 |
| Regular (PAYG) — Tier 3 | 1,000 |
| Regular (PAYG) — Tier 4 | 2,000 |
| Enterprise | 2,000 |
Unlike the per-minute limits, **Enterprise organizations are not exempt** — they get the elevated ceiling instead. Requests over the limit receive a retryable `429`:
```http
HTTP/1.1 429 Too Many Requests
Retry-After: 1
```
```json
{
"error": {
"message": "Too many concurrent requests for this organization (limit: 100). Retry shortly, or reduce request concurrency.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded"
}
}
```
Because slots free up continuously as in-flight requests complete, retrying after a short backoff typically succeeds — there is no fixed window to wait out. If you consistently hit the concurrency limit, reduce your client-side parallelism or [contact us](mailto:contact@vichar.io) about raising your ceiling.
## Free Models [#free-models]
Free models (models with zero input and output pricing) have additional rate limits that depend on your account's credit status. The limit is counted **per organization per model id**, so each free model has its own independent budget. Using free models also requires a **verified email address** — unverified accounts receive a `403`.
### Base Rate Limits [#base-rate-limits]
For organizations with **zero credits**:
* **5 requests per 10 minutes** per free model
* Resets every 10 minutes
### Elevated Rate Limits [#elevated-rate-limits]
For organizations that have **purchased at least some credits**:
* **20 requests per minute** per free model
* Resets every minute
When using free models with elevated limits, your credits will **not** be
deducted. The elevated rate limits are simply a benefit for users who have
added credits to their account.
## Provider Limits [#provider-limits]
Vichar uses configured provider/model caps to route requests away from providers that have reached their allowance. If you pin a provider and disable fallback, reaching one of these gateway-enforced caps can return `429`.
Rate limits returned by an upstream provider are treated as provider errors and are eligible for retries or [fallback routing](https://docs.vichar.io/features/routing), subject to your routing configuration and available providers. Upstream `Retry-After` and rate-limit headers are never forwarded to clients, including when retries are exhausted. Provider-scoped quota headers (`X-RateLimit-*-Provider*`) are not exposed.
## Rate Limit Headers [#rate-limit-headers]
All rate-limit response headers describe limits enforced by **Vichar**.
Successful authenticated responses carry the organization requests-per-minute (RPM) policy, remaining request quota, and reset delay only when they passed an RPM quota check. Anonymous requests, enterprise RPM exemptions, and disabled or unavailable RPM limiters omit these headers on success.
```http
RateLimit-Policy: "requests";q=600;w=60
RateLimit: "requests";r=599;t=60
RateLimit-Limit: 600
RateLimit-Remaining: 599
RateLimit-Reset: 60
```
`RateLimit-Policy` and `RateLimit` use the HTTP Structured Fields syntax in [draft-ietf-httpapi-ratelimit-headers-11](https://datatracker.ietf.org/doc/html/draft-ietf-httpapi-ratelimit-headers-11), currently an Internet-Draft. `q` is the quota, `w` the window in seconds, `r` the remaining quota, and `t` a reset delay in seconds. The request window rolls; the reset delay on an admitted request is a conservative upper bound, and quota may become available sooner.
Organization throttles return 429 with `Retry-After` in seconds, zero remaining quota, and a reset delay. Concurrency throttles identify the `"concurrency"` policy with `qu="concurrent-requests"` and suggest retrying after one second; capacity depends on requests finishing.
The legacy `RateLimit-Limit`, `RateLimit-Remaining`, and `RateLimit-Reset` fields remain available. `RateLimit-Reset` is a delay in seconds; `X-RateLimit-Reset` is a Unix timestamp. These fields also remain available as `X-RateLimit-Limit` and `X-RateLimit-Remaining`. Browser clients on allowed CORS origins can read these headers.
Other gateway gates, such as spending limits, can reject requests independently of the advertised organization quota. Honor `Retry-After` when supplied; otherwise use exponential backoff with jitter. Inspect the error body to distinguish a temporary throttle from a spending limit that needs account action.
## Rate Limit Exceeded [#rate-limit-exceeded]
When you exceed a rate limit, you'll receive a `429 Too Many Requests` response:
```json
{
"error": {
"message": "Rate limit exceeded for /v1/chat/completions. Please retry after 12 seconds.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded"
}
}
```
This uses the standard OpenAI-compatible error envelope. Requests to the Anthropic-compatible `/v1/messages` endpoint receive the Anthropic error shape instead. See [Error Handling](https://docs.vichar.io/resources/error-handling) for the full format and status-code reference.
## Gateway Overload (529) [#gateway-overload-529]
Separately from per-account rate limits, the gateway protects itself from
transient overload. When a single gateway instance is holding too many
concurrent in-flight inference requests at once — across all organizations
combined (for example during a traffic spike, or when an upstream provider is
slow and connections pile up) — it sheds excess inference requests with an
`HTTP 529` response instead of letting them queue indefinitely. Non-inference
endpoints such as the models list are unaffected:
```http
HTTP/1.1 529
Retry-After: 1
```
```json
{
"error": {
"message": "Gateway overloaded, please retry",
"type": "overloaded",
"code": "overloaded"
}
}
```
Requests to the Anthropic-compatible `/v1/messages` endpoint receive the
equivalent Anthropic envelope (`{ "type": "error", "error": { "type": "overloaded_error" } }`),
matching Anthropic's own `529` behavior.
A `529` is **transient and retryable** — it reflects momentary capacity, not a
quota on your account. Unlike a `429`, it is not tied to your credits or model
tier, and retrying after a short delay (honoring the `Retry-After` header)
will typically succeed.
How `529` differs from `429`:
| | `429 Too Many Requests` | `529` Overloaded |
| --------- | ---------------------------------------------------------------------------------------------- | -------------------------------------- |
| Cause | Your organization exceeded its request rate or [concurrency](#concurrent-request-limits) limit | The gateway is momentarily at capacity |
| Scope | Per organization / API key | Transient, server-side |
| Fix | Slow down or reduce concurrency; add credits for [elevated limits](#elevated-rate-limits) | Retry after a short delay |
| Retryable | After the window resets (rate) or as soon as an in-flight request finishes (concurrency) | Yes, immediately with backoff |
## Best Practices [#best-practices]
* **Respect `Retry-After`.** Implement exponential backoff when you receive `429` or `529` responses, starting from the `Retry-After` value.
* **Watch the headers.** Monitor `RateLimit` or `RateLimit-Remaining` to back off before you hit the limit.
* **Spread traffic across endpoints.** Limits are per endpoint, so unrelated workloads don't compete for the same budget.
* **Scale with usage.** Regular organizations unlock higher limits automatically as lifetime spend grows; [contact us](mailto:contact@vichar.io) about an [Enterprise plan](mailto:contact@vichar.io) to remove the per-minute limits and get an elevated concurrency ceiling.
Adding even a small amount of credits to your account (e.g., $20) will
immediately upgrade your free model rate limits from 5 requests per 10 minutes
to 20 requests per minute (free-model use still requires a verified email).
# AWS
URL: https://docs.vichar.io/self-host/aws
This guide covers a production deployment of Vichar on AWS using EKS for the application services and managed AWS services for the backing stores.
## Architecture [#architecture]
| Component | AWS service |
| ---------- | ---------------------------- |
| Compute | Amazon EKS (Kubernetes) |
| PostgreSQL | Amazon RDS for PostgreSQL |
| Redis | Amazon ElastiCache for Redis |
| Secrets | AWS Secrets Manager |
| Ingress | AWS Load Balancer Controller |
## What to configure [#what-to-configure]
### 1. PostgreSQL — Amazon RDS [#1-postgresql--amazon-rds]
Create an RDS for PostgreSQL instance in a private subnet. Enable automated backups and Multi-AZ for high availability. Note the connection string for `DATABASE_URL`.
### 2. Redis — Amazon ElastiCache [#2-redis--amazon-elasticache]
Create an ElastiCache for Redis cluster in the same VPC. Use it for `REDIS_URL`. A single-node cluster is fine to start; enable replication for production.
### 3. Compute — Amazon EKS [#3-compute--amazon-eks]
Create an EKS cluster with a managed node group sized for your traffic. Install the [AWS Load Balancer Controller](https://kubernetes-sigs.github.io/aws-load-balancer-controller/) to expose the gateway through an Application Load Balancer.
### 4. Networking [#4-networking]
Run RDS and ElastiCache in private subnets and allow inbound traffic only from the EKS node security group. Expose only the gateway (and optionally the UI) to the internet through the load balancer.
### 5. Secrets — AWS Secrets Manager [#5-secrets--aws-secrets-manager]
Store `AUTH_SECRET`, `GATEWAY_API_KEY_HASH_SECRET`, and your provider API keys in Secrets Manager. Sync them into the cluster with the [External Secrets Operator](https://external-secrets.io/) or the AWS Secrets and Configuration Provider (ASCP).
## Deploy the Helm chart [#deploy-the-helm-chart]
With the backing services in place, deploy Vichar with the [Helm chart](https://docs.vichar.io/self-host/kubernetes), pointing it at your RDS and ElastiCache endpoints:
```bash
helm install llmgateway oci://ghcr.io/theopenco/charts/llmgateway -f values.yaml
```
```yaml
config:
DATABASE_URL: "postgres://user:password@your-rds-endpoint:5432/llmgateway"
REDIS_URL: "redis://your-elasticache-endpoint:6379"
AUTH_SECRET: "from-secrets-manager"
GATEWAY_API_KEY_HASH_SECRET: "from-secrets-manager"
```
See the [Kubernetes guide](https://docs.vichar.io/self-host/kubernetes) for the full set of configurable values and how to scale the gateway.
# Azure
URL: https://docs.vichar.io/self-host/azure
This guide covers a production deployment of Vichar on Azure using AKS for the application services and managed Azure services for the backing stores.
## Architecture [#architecture]
| Component | Azure service |
| ---------- | --------------------------------- |
| Compute | Azure Kubernetes Service (AKS) |
| PostgreSQL | Azure Database for PostgreSQL |
| Redis | Azure Cache for Redis |
| Secrets | Azure Key Vault |
| Ingress | Application Gateway / AKS Ingress |
## What to configure [#what-to-configure]
### 1. PostgreSQL — Azure Database for PostgreSQL [#1-postgresql--azure-database-for-postgresql]
Create an Azure Database for PostgreSQL Flexible Server with private access. Enable automated backups and zone-redundant high availability. Note the connection details for `DATABASE_URL`.
### 2. Redis — Azure Cache for Redis [#2-redis--azure-cache-for-redis]
Create an Azure Cache for Redis instance in the same virtual network and use its endpoint for `REDIS_URL`. Choose a Standard or Premium tier for replication in production.
### 3. Compute — AKS [#3-compute--aks]
Create an AKS cluster with a node pool sized for your traffic. Use the [Application Gateway Ingress Controller](https://learn.microsoft.com/en-us/azure/application-gateway/ingress-controller-overview) or an NGINX ingress to expose the gateway.
### 4. Networking [#4-networking]
Deploy the database and cache with private endpoints inside your virtual network and restrict access to the AKS subnet. Expose only the gateway (and optionally the UI) to the internet.
### 5. Secrets — Azure Key Vault [#5-secrets--azure-key-vault]
Store `AUTH_SECRET`, `GATEWAY_API_KEY_HASH_SECRET`, and your provider API keys in Key Vault. Sync them into the cluster with the [Azure Key Vault Provider for Secrets Store CSI Driver](https://learn.microsoft.com/en-us/azure/aks/csi-secrets-store-driver).
## Deploy the Helm chart [#deploy-the-helm-chart]
With the backing services in place, deploy Vichar with the [Helm chart](https://docs.vichar.io/self-host/kubernetes), pointing it at your Azure Database and Azure Cache endpoints:
```bash
helm install llmgateway oci://ghcr.io/theopenco/charts/llmgateway -f values.yaml
```
```yaml
config:
DATABASE_URL: "postgres://user:password@your-postgres-host:5432/llmgateway"
REDIS_URL: "redis://your-redis-host:6380"
AUTH_SECRET: "from-key-vault"
GATEWAY_API_KEY_HASH_SECRET: "from-key-vault"
```
See the [Kubernetes guide](https://docs.vichar.io/self-host/kubernetes) for the full set of configurable values and how to scale the gateway.
# Docker Compose
URL: https://docs.vichar.io/self-host/docker-compose
Docker Compose runs each service in its own container, giving you more control over scaling and configuration than the [single Docker image](https://docs.vichar.io/self-host/docker). It's a good fit for a single production host. For multi-node, high-availability deployments, use [Kubernetes](https://docs.vichar.io/self-host/kubernetes).
## Prerequisites [#prerequisites]
* Latest Docker with Compose
* API keys for the LLM providers you want to use (OpenAI, Anthropic, etc.)
## Option A: Unified image with Compose [#option-a-unified-image-with-compose]
Run the all-in-one image under Compose for easy lifecycle management:
```bash
# Download the compose file
curl -O https://raw.githubusercontent.com/theopenco/llmgateway/main/infra/docker-compose.unified.yml
curl -O https://raw.githubusercontent.com/theopenco/llmgateway/main/.env.unified.example
# Configure environment
cp .env.unified.example .env
# Edit .env with your configuration
# Start the service
docker compose -f docker-compose.unified.yml up -d
```
## Option B: Separate services [#option-b-separate-services]
Run each service in its own container for the most flexibility:
```bash
# Clone the repository
git clone https://github.com/vicharai/api.git
cd llmgateway
# Configure environment
cp .env.example .env
# Edit .env with your configuration
# Start the services
docker compose -f infra/docker-compose.split.yml up -d
```
Pin to a specific version by replacing the `latest` image tags with the [latest release](https://github.com/vicharai/api/releases). To build the images from source, use the `*.local.yml` compose files in the `infra` directory.
## Accessing your instance [#accessing-your-instance]
* **Web Interface**: [http://localhost:3002](http://localhost:3002)
* **Chat Playground**: [http://localhost:3003](http://localhost:3003)
* **Documentation**: [http://localhost:3005](http://localhost:3005)
* **Provider Portal (Airside)**: [http://localhost:3007](http://localhost:3007)
* **API Endpoint**: [http://localhost:4002](http://localhost:4002)
* **Gateway Endpoint**: [http://localhost:4001](http://localhost:4001)
## Required configuration [#required-configuration]
At minimum, set these environment variables:
```bash
# Database (change the password!)
POSTGRES_PASSWORD=your_secure_password_here
# Authentication
AUTH_SECRET=your-secret-key-here
GATEWAY_API_KEY_HASH_SECRET=your-api-key-hash-secret-here
# LLM Provider API Keys (add the ones you need)
LLM_OPENAI_API_KEY=sk-...
LLM_ANTHROPIC_API_KEY=sk-ant-...
```
## Management commands [#management-commands]
```bash
# View logs
docker compose -f infra/docker-compose.split.yml logs -f
# Restart services
docker compose -f infra/docker-compose.split.yml restart
# Stop services
docker compose -f infra/docker-compose.split.yml down
```
## Managing provider credentials from the admin dashboard [#managing-provider-credentials-from-the-admin-dashboard]
Provider credentials can also live in the database instead of the environment.
The admin dashboard's **Provider Credentials** page lets you add, edit and remove
them without redeploying, and each credential carries every setting the provider
needs — API key, base URL, project, region, resource, API version — plus a free-form
note so several keys for the same provider stay tellable apart.
Resolution is per provider: as soon as a provider has at least one managed
credential, credits-mode requests to it use those credentials and its `LLM_*`
environment variables are ignored entirely. Providers with no managed credential
keep reading the environment, so you can migrate one provider at a time.
Managed credentials cover the same axes as the environment variables they replace:
* Several credentials per provider, selected with the same health-aware routing
as the comma-separated env lists.
* A per-credential audience (all organizations / enterprise plans), mirroring
the `__ENTERPRISE` and `__PLANS` suffixes below.
* An optional region, mirroring the `{ENV_VAR}__{REGION}` overrides.
A request that resolves no region is served from the provider's default region,
so it is served by a region-agnostic credential or by one pinned to that default
region. Providers whose keys are region-scoped — those with no global region, so
every credential belongs to exactly one region — are therefore fully covered by
one credential per region, with no region-agnostic credential needed. Providers
whose key works across every region (AWS Bedrock) are stricter: a region on the
credential scopes it to that region only, and a request for another region fails
rather than borrowing it.
Video generation pins the credential that created a job onto the job itself, so
polling and content retrieval hours later go back through the same credential
rather than re-selecting one.
Provider credential tokens are encrypted at rest with AES-256-GCM because the
gateway must recover them for upstream requests. New and rolled gateway API
keys are stored only as keyed HMAC-SHA-256 fingerprints; incoming secrets are
fingerprinted and compared during authentication. Both protections derive from
`GATEWAY_API_KEY_HASH_SECRET`, so there is no extra variable to set. Give it a
strong random value (`openssl rand -base64 32`) and configure that same value
on every service.
For rotation, prepend the same new secret to the comma-separated keyring on
every service and retain the old entries while credentials and hashes still use
them. Authentication checks every retained entry, while new hashes use the
first entry. Removing an old entry makes its provider credentials undecryptable
and its remaining hashes unusable.
## Multiple API keys and load balancing [#multiple-api-keys-and-load-balancing]
Vichar supports multiple API keys per provider for load balancing and increased availability. Provide comma-separated values:
```bash
# Multiple OpenAI keys for load balancing
LLM_OPENAI_API_KEY=sk-key1,sk-key2,sk-key3
# Multiple Anthropic keys
LLM_ANTHROPIC_API_KEY=sk-ant-key1,sk-ant-key2
```
### Health-aware routing [#health-aware-routing]
The gateway tracks the health of each API key and routes requests to healthy keys. If a key returns consecutive errors, it's temporarily skipped. Keys that return authentication errors (401/403) are blacklisted until restart.
### Related configuration values [#related-configuration-values]
For providers that require additional configuration (like Google Vertex), specify multiple values that correspond to each API key. The gateway uses the matching index:
```bash
# Multiple Google Vertex configurations
LLM_GOOGLE_VERTEX_API_KEY=key1,key2,key3
LLM_GOOGLE_CLOUD_PROJECT=project-a,project-b,project-c
LLM_GOOGLE_VERTEX_REGION=us-central1,europe-west1,asia-east1
```
When the gateway selects `key2`, it automatically uses `project-b` and `europe-west1`. If you have fewer configuration values than keys, the last value is reused for the remaining keys.
`LLM_GOOGLE_CLOUD_PROJECT` is optional for Google Vertex API-key chat, embedding, and speech requests, which use the projectless publisher-model endpoint when it is unset. Set it for OAuth authentication, video generation, or project-scoped Vertex URLs.
## Enterprise and plan provider env overrides [#enterprise-and-plan-provider-env-overrides]
Any provider env var — API keys, base URLs, regions, Google Cloud projects, Azure resources, and other provider-specific settings — supports optional per-audience overrides:
* `__ENTERPRISE` suffix: used instead of the base var for organizations on the enterprise plan
* `__PLANS` suffix: used instead of the base var for plan-based (non-PAYG) organizations
```bash
# Shared key for all organizations
LLM_OPENAI_API_KEY=sk-shared-key
# Used instead of the shared key, but only for enterprise-plan organizations
LLM_OPENAI_API_KEY__ENTERPRISE=sk-enterprise-key
# Used instead of the shared key, but only for plan-based organizations
LLM_OPENAI_API_KEY__PLANS=sk-plans-key
# Companion settings can be overridden the same way, e.g. a dedicated
# Google Cloud project and base URL per audience
LLM_GOOGLE_CLOUD_PROJECT__ENTERPRISE=enterprise-gcp-project
LLM_OPENAI_BASE_URL__PLANS=https://plans-proxy.internal
```
Notes:
* Overrides are optional — matching organizations fall back to the base var when an override is unset. All other organizations never read the override vars.
* The base API key is still what makes a provider available for routing, so set it even when you route all enterprise or plan traffic through dedicated keys.
* Comma-separated values, load balancing, and health-aware routing work exactly like the base key; each override key list gets its own independent health tracking.
* A set override replaces the base var wholesale, including its comma-separated list. Indexed companion values (like `LLM_GOOGLE_CLOUD_PROJECT`) resolve at the selected key index from the override list when one is set, otherwise from the base list — so keep an override key list index-compatible with its companion lists (or use single values).
* Region-specific overrides compose with the variant suffix. For matching organizations the API-key lookup order is `{BASE}__ENTERPRISE__{REGION}` → `{BASE}__{REGION}` → `{BASE}__ENTERPRISE` → `{BASE}` (same pattern with `__PLANS`).
* If an organization is both on the enterprise plan and a plan-based org, the enterprise overrides win.
# Docker
URL: https://docs.vichar.io/self-host/docker
The unified Docker image bundles every service — UI, API, Gateway, PostgreSQL, and Redis — into a single container. It's the fastest way to get a working instance and is ideal for trying Vichar out or running a single low-traffic deployment.
For production, run each service separately with [Docker Compose](https://docs.vichar.io/self-host/docker-compose) or deploy to [Kubernetes](https://docs.vichar.io/self-host/kubernetes) with managed Postgres and Redis.
## Prerequisites [#prerequisites]
* Latest Docker
* API keys for the LLM providers you want to use (OpenAI, Anthropic, etc.)
## Run the container [#run-the-container]
```bash
# Set a strong secret first
export LLM_GATEWAY_SECRET="your-secret-key-here"
export GATEWAY_API_KEY_HASH_SECRET="your-api-key-hash-secret-here"
# Required only for licensed Enterprise features
export LLMGATEWAY_ENTERPRISE_LICENSE="your-signed-license"
# Run the container
docker run -d \
--name llmgateway \
--restart unless-stopped \
-p 3002:3002 \
-p 3003:3003 \
-p 3005:3005 \
-p 3006:3006 \
-p 3007:3007 \
-p 4001:4001 \
-p 4002:4002 \
-v llmgateway_postgres:/var/lib/postgresql/data \
-v llmgateway_redis:/var/lib/redis \
-e AUTH_SECRET="$LLM_GATEWAY_SECRET" \
-e GATEWAY_API_KEY_HASH_SECRET="$GATEWAY_API_KEY_HASH_SECRET" \
-e LLMGATEWAY_ENTERPRISE_LICENSE="$LLMGATEWAY_ENTERPRISE_LICENSE" \
ghcr.io/theopenco/llmgateway-unified:latest
```
Docker creates the named volumes automatically on first run. Do not bind-mount a host directory directly to `/var/lib/postgresql/data`, because PostgreSQL initialization inside the container needs to manage permissions on that path.
Pin to a specific version instead of `latest` using the [latest release tag](https://github.com/vicharai/api/releases).
## Accessing your instance [#accessing-your-instance]
* **Web Interface**: [http://localhost:3002](http://localhost:3002)
* **Chat Playground**: [http://localhost:3003](http://localhost:3003)
* **Documentation**: [http://localhost:3005](http://localhost:3005)
* **Provider Portal (Airside)**: [http://localhost:3007](http://localhost:3007)
* **API Endpoint**: [http://localhost:4002](http://localhost:4002)
* **Gateway Endpoint**: [http://localhost:4001](http://localhost:4001)
## Management commands [#management-commands]
```bash
# View logs
docker logs llmgateway
# Restart the container
docker restart llmgateway
# Stop the container
docker stop llmgateway
```
## Required configuration [#required-configuration]
At minimum, set these environment variables:
```bash
# Authentication
AUTH_SECRET=your-secret-key-here
GATEWAY_API_KEY_HASH_SECRET=your-api-key-hash-secret-here
# LLM Provider API Keys (add the ones you need)
LLM_OPENAI_API_KEY=sk-...
LLM_ANTHROPIC_API_KEY=sk-ant-...
```
## Next steps [#next-steps]
Once your instance is running:
1. **Open the web interface** at [http://localhost:3002](http://localhost:3002)
2. **Create your first organization** and project
3. **Generate API keys** for your applications
4. **Test the gateway** by making API calls to [http://localhost:4001](http://localhost:4001)
# Google Cloud
URL: https://docs.vichar.io/self-host/gcp
This guide covers a production deployment of Vichar on Google Cloud using GKE for the application services and managed Google Cloud services for the backing stores.
## Architecture [#architecture]
| Component | Google Cloud service |
| ---------- | ---------------------------------- |
| Compute | Google Kubernetes Engine (GKE) |
| PostgreSQL | Cloud SQL for PostgreSQL |
| Redis | Memorystore for Redis |
| Secrets | Secret Manager |
| Ingress | GKE Ingress / Cloud Load Balancing |
## What to configure [#what-to-configure]
### 1. PostgreSQL — Cloud SQL [#1-postgresql--cloud-sql]
Create a Cloud SQL for PostgreSQL instance with a private IP. Enable automated backups and high availability. Note the connection details for `DATABASE_URL`.
### 2. Redis — Memorystore [#2-redis--memorystore]
Create a Memorystore for Redis instance in the same VPC and use its endpoint for `REDIS_URL`. Enable a read replica for production.
### 3. Compute — GKE [#3-compute--gke]
Create a GKE cluster (Autopilot or Standard) sized for your traffic. GKE provisions an HTTP(S) load balancer automatically when you create an Ingress for the gateway.
### 4. Networking [#4-networking]
Use a private VPC and connect Cloud SQL via Private Service Access and Memorystore via its private endpoint. Restrict access so only the GKE workloads can reach the database and cache, and expose only the gateway to the internet.
### 5. Secrets — Secret Manager [#5-secrets--secret-manager]
Store `AUTH_SECRET`, `GATEWAY_API_KEY_HASH_SECRET`, and your provider API keys in Secret Manager. Sync them into the cluster with the [External Secrets Operator](https://external-secrets.io/) or the Secret Manager CSI driver.
## Deploy the Helm chart [#deploy-the-helm-chart]
With the backing services in place, deploy Vichar with the [Helm chart](https://docs.vichar.io/self-host/kubernetes), pointing it at your Cloud SQL and Memorystore endpoints:
```bash
helm install llmgateway oci://ghcr.io/theopenco/charts/llmgateway -f values.yaml
```
```yaml
config:
DATABASE_URL: "postgres://user:password@your-cloud-sql-ip:5432/llmgateway"
REDIS_URL: "redis://your-memorystore-ip:6379"
AUTH_SECRET: "from-secret-manager"
GATEWAY_API_KEY_HASH_SECRET: "from-secret-manager"
```
See the [Kubernetes guide](https://docs.vichar.io/self-host/kubernetes) for the full set of configurable values and how to scale the gateway.
# Self Host Vichar
URL: https://docs.vichar.io/self-host
Vichar is a self-hostable platform that provides a unified API gateway for multiple LLM providers. Run it on your own infrastructure to keep full control over your data and avoid platform fees.
Pick the deployment path that matches where you're running it.
Self-hosting guides are documented under https://docs.vichar.io/self-host and included in full in this file.
## Which option should I choose? [#which-option-should-i-choose]
* **Trying it out or running a single low-traffic instance?** Start with [Docker](https://docs.vichar.io/self-host/docker) or [Docker Compose](https://docs.vichar.io/self-host/docker-compose) on one machine.
* **Running in production?** Deploy to [Kubernetes](https://docs.vichar.io/self-host/kubernetes) with our Helm chart, and use a managed Postgres and Redis from your cloud.
* **On a specific cloud?** Follow the [AWS](https://docs.vichar.io/self-host/aws), [Google Cloud](https://docs.vichar.io/self-host/gcp), or [Azure](https://docs.vichar.io/self-host/azure) guide for the exact managed services to provision.
## What you'll need [#what-youll-need]
Every deployment is built from the same pieces:
* **Stateless services** — the gateway, API, UI, and a background worker. Scale these freely; they hold no data between requests.
* **PostgreSQL** — the source of truth for users, projects, keys, and usage records.
* **Redis** — response caching and the queue that feeds the worker.
* **Provider API keys** — the OpenAI, Anthropic, Google, and other credentials the gateway uses, injected as secrets.
In production, run PostgreSQL and Redis as managed services from your cloud provider so backups, failover, and patching are handled for you.
## Product links [#product-links]
Set `UI_URL` to your Vichar dashboard origin. When unset, it uses `https://app.vichar.io`; with `NODE_ENV=development`, it uses `http://localhost:3002`.
## Sign-up email restrictions [#sign-up-email-restrictions]
With `HOSTED=true`, administrators can manage the custom email-domain blocklist in
**Admin → Settings → Blocked sign-up email domains**. Paste domains one per line
or separated by commas, then save. Entries also block subdomains; updates apply
to subsequent email sign-ups without a restart. Existing users can still sign in.
The custom list starts empty and has no built-in entries. Clearing and saving it
removes custom domain blocks. The separate disposable-email and plus-address
checks remain active in hosted mode.
# Kubernetes
URL: https://docs.vichar.io/self-host/kubernetes
Kubernetes is the deployment model we recommend for production. The gateway is stateless, so it scales horizontally with no coordination; pods self-heal; and the same manifests run on any cluster — EKS, GKE, AKS, or your own.
We publish an official **Helm chart** that deploys the gateway, API, UI, and worker with sane defaults and lets you wire in a managed Postgres and Redis through values.
## Prerequisites [#prerequisites]
* A Kubernetes cluster and `kubectl` configured to reach it
* [Helm 3](https://helm.sh/docs/intro/install/)
* A PostgreSQL database and Redis instance (use a managed service in production)
## Install the chart [#install-the-chart]
The chart is published as an OCI artifact on GitHub Container Registry:
```bash
helm install llmgateway oci://ghcr.io/theopenco/charts/llmgateway
```
This installs the latest published version. To pin to a specific release, append `--version `, matching a published release tag without the `v` prefix (e.g. `1.2.3`).
## Configure with values [#configure-with-values]
Provide a `values.yaml` to point the chart at your managed database and cache and to set your secrets:
```yaml
config:
AUTH_SECRET: "your-secret-key-here"
GATEWAY_API_KEY_HASH_SECRET: "your-api-key-hash-secret-here"
DATABASE_URL: "postgres://user:password@your-managed-host:5432/llmgateway"
REDIS_URL: "redis://your-managed-host:6379"
LLM_OPENAI_API_KEY: "sk-..."
LLM_ANTHROPIC_API_KEY: "sk-ant-..."
```
```bash
helm install llmgateway oci://ghcr.io/theopenco/charts/llmgateway -f values.yaml
```
## Scaling the gateway [#scaling-the-gateway]
Because the gateway is stateless, scale it horizontally to match traffic — either by setting the replica count in values or with a `HorizontalPodAutoscaler` that targets CPU or request load. The API, UI, and worker serve your team rather than your traffic and rarely need more than one or two replicas.
See the [Helm chart README](https://github.com/vicharai/api/tree/main/infra/helm) for the full list of configurable values and the [list of available versions](https://github.com/vicharai/api/pkgs/container/charts%2Fllmgateway).
# Gateway Caching
URL: https://docs.vichar.io/features/caching/gateway-caching
Gateway caching serves a previously-seen request entirely from Vichar without forwarding it to the upstream provider. Requests are matched on the parsed [cache-key fields](#cache-key-generation) — JSON formatting outside string values, and parameters outside the cache key, do not create a separate entry; the text of your messages and the key order of message and tool objects do. Repeated identical calls cost **$0** — there is no inference and no provider charge. It is most useful for API workloads with deterministic inputs (classification, batch jobs, FAQ lookups, retries) rather than free-form chat.
If you want to reduce the cost of long, partially-shared prompts in chat apps
or coding tools, you want [Provider Cache
Control](https://docs.vichar.io/features/caching/provider-cache-control) instead. That discounts the
cached portion of your prompt on every call — it does not require identical
requests. See the [Caching Overview](https://docs.vichar.io/features/caching) for a side-by-side
comparison.
## How It Works [#how-it-works]
When you make an API request:
1. Vichar generates a cache key based on the request parameters
2. If a matching cached response exists, it's returned immediately
3. If no cache exists, the request is forwarded to the provider
4. The response is cached for future identical requests
This means repeated identical requests are served instantly from cache without incurring additional provider costs.
## Cost Savings [#cost-savings]
Caching can dramatically reduce costs for applications with repetitive requests:
| Scenario | Without Caching | With Caching | Savings |
| --------------------------- | --------------- | ------------ | ------- |
| 1,000 identical requests | $10.00 | $0.01 | 99.9% |
| 50% duplicate rate | $10.00 | $5.00 | 50% |
| Retry after transient error | $0.02 | $0.01 | 50% |
Cached responses are free from provider costs. You only pay for the initial
request that populates the cache.
## Requirements [#requirements]
Caching is **free** and **independent** of [Data
Retention](https://docs.vichar.io/features/data-retention). Cached responses live in a short-lived
cache bounded by your configured TTL (60 seconds by default) and are not
stored as long-term request data — you do not need to enable data retention to
use caching.
To use caching:
1. Enable **Caching** in your project settings under Preferences
2. Configure the cache duration (TTL) as needed
3. Make requests as normal—caching is automatic
Gateway caching is not available while a zero data retention policy is active
— the project setting is ignored there and every request goes upstream.
## Cache Key Generation [#cache-key-generation]
The cache key is scoped to your **project** and includes the **resolved provider and model** — so two projects never share cache entries, and the same request routed to a different provider is a separate entry. The key is a hash of these request parameters:
* Resolved provider and model
* Messages array (roles and content, including the system prompt)
* Temperature, max tokens, top P
* Frequency and presence penalty
* Response format
* Tools/functions, tool choice, and the web search tool
* Reasoning effort and reasoning max tokens
* `prompt_cache_key`, `prompt_cache_retention`, `prompt_cache_options`
* `n` and `service_tier`
* The response mode — streaming and non-streaming entries are kept separate, so an otherwise identical `stream: true` request never shares an entry with a non-streaming one
Requests with different values for any of these parameters, even slight
variations, will not share cache entries. Parameters outside this list do not
affect the cache key.
## Cache Behavior [#cache-behavior]
### Cache Hits [#cache-hits]
When a cache hit occurs:
* Non-streaming responses are returned immediately
* No provider API call is made
* No inference costs are incurred
### Cache Misses [#cache-misses]
When a cache miss occurs:
* Request is forwarded to the LLM provider
* A successful response is stored in cache — errored, client-cancelled, or empty responses are never stored (a response truncated by `max_tokens` is cached like any other)
* Normal inference costs apply
* Future identical requests will hit the cache
## Streaming and Caching [#streaming-and-caching]
Caching works with both streaming and non-streaming requests:
* **Non-streaming**: Full response is cached and returned immediately on a hit
* **Streaming**: The complete stream is cached chunk by chunk (only once it finished successfully) and replayed on a hit, reproducing the original chunk timing with each gap capped at one second — so a streamed replay takes roughly as long as the original stream
## Cache TTL (Time-to-Live) [#cache-ttl-time-to-live]
Cache duration is configurable per project in your project settings. You can set the cache TTL from 10 seconds up to 1 year (31,536,000 seconds).
The default cache duration is 60 seconds. Adjust this based on your use case—longer durations work well for static content, while shorter durations are better for frequently changing data.
## Identifying Cached Responses [#identifying-cached-responses]
A cache hit replays the stored completion — same `id`, same content, same token counts. The `metadata` envelope is rebuilt for the current request (fresh `log_id`, and a fresh `request_id` on non-streaming responses) and all cost fields are zeroed, so the body is not byte-for-byte identical to the original response. Two markers tell you it was a replay:
* the `x-llmgateway-cache: HIT` response header — set on `/v1/chat/completions`, `/v1/messages` (which has no metadata envelope), and streaming AI SDK requests; it is not surfaced on `/v1/responses` or on non-streaming AI SDK responses
* `metadata.cached: true` on the response body (and on the final metadata chunk of a streamed replay)
All cost fields are zeroed on a replay, because no upstream call was made:
```json
{
"usage": {
"prompt_tokens": 12,
"completion_tokens": 48,
"total_tokens": 60,
"cost": 0,
"cost_details": {
"total_cost": 0,
"input_cost": 0,
"output_cost": 0
}
},
"metadata": {
"cached": true
}
}
```
Token counts are **not** zeroed — they still describe the completion you are receiving, and are what the dashboard records for analytics. Note that any prompt-caching fields (`prompt_tokens_details.cached_tokens`, `cache_write_tokens`) describe the *original* upstream call, not the replay.
## Bypassing the Cache for a Single Request [#bypassing-the-cache-for-a-single-request]
Send `x-no-cache: true` to skip the cache for one request — the call goes upstream and its response is not stored. Useful when a client retries an identical request and expects a fresh sample (for example an agent loop) without disabling caching for the whole project. The header works on `/v1/chat/completions`, `/v1/messages`, and the AI SDK endpoints; it is not forwarded on `/v1/responses`.
```bash
curl https://api.vichar.io/v1/chat/completions \
-H "Authorization: Bearer $LLM_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-H "x-no-cache: true" \
-d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Hello"}]}'
```
## Use Cases [#use-cases]
### Development and Testing [#development-and-testing]
During development, you often send the same prompts repeatedly:
```typescript
// This prompt will only incur costs once
const response = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "Explain quantum computing" }],
});
```
### Chatbots with Common Questions [#chatbots-with-common-questions]
FAQ-style interactions often have repeated questions:
```typescript
// Common questions are served from cache
const faqs = [
"What are your business hours?",
"How do I reset my password?",
"What is your return policy?",
];
```
### Batch Processing [#batch-processing]
Processing large datasets with potentially duplicate items:
```typescript
// Duplicate items in batch are served from cache
for (const item of items) {
const response = await client.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: `Classify: ${item}` }],
});
}
```
## Best Practices [#best-practices]
### Maximize Cache Hits [#maximize-cache-hits]
* Use consistent prompt formatting
* Normalize input data before sending
* Use deterministic parameters (temperature: 0)
* Avoid including timestamps or random values in prompts
### Appropriate Use Cases [#appropriate-use-cases]
Caching is most effective for:
* Static knowledge queries
* Classification tasks
* FAQ responses
* Development/testing
* Retry scenarios
### When to Avoid Caching [#when-to-avoid-caching]
Caching may not be suitable for:
* Real-time data requirements
* Highly personalized responses
* Time-sensitive information
* Creative tasks requiring variety
* Chat or coding tools where prompts overlap but are not identical — use [Provider Cache Control](https://docs.vichar.io/features/caching/provider-cache-control) instead
## Pricing [#pricing]
Caching is **completely free**. Cached responses are held in a short-lived
Redis-backed cache (bounded by your configured TTL) and do not incur storage
charges. Storage costs only apply if you separately enable [Data
Retention](https://docs.vichar.io/features/data-retention) for full request/response payloads.
Caching reduces both inference cost and latency at no additional charge.
# Caching
URL: https://docs.vichar.io/features/caching
Vichar supports **two distinct kinds of caching**, and they solve different problems. Pick the one that matches your workload — they can also be used together.
## Provider / Model Caching [#provider--model-caching]
The provider performs the caching. When your request reuses a long prefix from a previous call (a system prompt, conversation history, tool definitions, a long document), the model serves that prefix from its prompt cache and bills it at a reduced rate. New input tokens and **all output tokens are still billed at the normal rate** — only the cached portion is discounted.
This is the type of caching that powers efficient chat-based and assistant-based interactions, including chat apps and coding tools (Cursor, Cline, Claude Code, etc.) where the same context is reused turn after turn.
You see it in your usage as `prompt_tokens_details.cached_tokens`. For most providers it works automatically; some (notably Anthropic) also let you mark blocks explicitly with `cache_control` and choose a longer TTL.
The admin model stats show **Provider cache rate** beside cached token totals and in model and per-provider history charts. It is provider-cached input tokens divided by total input tokens (which already include cached tokens), multiplied by 100. Window summaries use summed token counts; output tokens and gateway cache hits are excluded. A dash means there are no input tokens in the selected window.
→ **[Read the Provider Cache Control docs](https://docs.vichar.io/features/caching/provider-cache-control)**
## Gateway Caching [#gateway-caching]
Vichar performs the caching. When a request is **identical** to a previous one (same model, same messages, same parameters), the response is served from the gateway's cache without any provider call. Repeated identical calls cost **$0**.
This is most useful for deterministic API workloads — classification, batch jobs, FAQ lookups, retries — rather than free-form chat, because chat prompts almost always differ on the latest turn.
→ **[Read the Gateway Caching docs](https://docs.vichar.io/features/caching/gateway-caching)**
## Which one do I want? [#which-one-do-i-want]
| If you… | Use |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| Build a chat app, assistant, or coding tool | [Provider Cache Control](https://docs.vichar.io/features/caching/provider-cache-control) |
| Send long system prompts or growing conversation history | [Provider Cache Control](https://docs.vichar.io/features/caching/provider-cache-control) |
| Want longer cache lifetimes than the provider default | [Provider Cache Control](https://docs.vichar.io/features/caching/provider-cache-control) (explicit `cache_control`) |
| Send the exact same request many times (batches, retries, FAQs) | [Gateway Caching](https://docs.vichar.io/features/caching/gateway-caching) |
| Want $0 on repeated calls instead of a discount | [Gateway Caching](https://docs.vichar.io/features/caching/gateway-caching) |
The two are not mutually exclusive. A coding tool can rely on provider caching
for its long system prompt **and** enable gateway caching so that
deterministic tool calls (e.g., file lookups) cost nothing on retry.
# Provider Cache Control
URL: https://docs.vichar.io/features/caching/provider-cache-control
Most modern LLM providers offer **prompt caching**: when a request reuses a long prefix from a previous request (for example, a multi-thousand-token system prompt or a growing conversation history), the provider stores that prefix and serves it back at a steep discount on subsequent calls. Only the cached portion is discounted — new input tokens and all output tokens are still billed at the normal rate.
This is the behavior you see surfaced as `cached_tokens` in your usage payloads, and it is what makes chat apps, assistants, and coding tools (Cursor, Cline, Claude Code, etc.) economically viable on long contexts.
Looking for $0 on repeated calls instead of a discount on the cached portion?
That is [Gateway Caching](https://docs.vichar.io/features/caching/gateway-caching), which serves
identical requests entirely from Vichar without hitting the provider. It is a
better fit for deterministic API workloads than for chat. See the [Caching
Overview](https://docs.vichar.io/features/caching) for a side-by-side comparison.
## Automatic caching [#automatic-caching]
For most users, prompt caching just works — you do not need to change your request payloads.
Providers including OpenAI, Anthropic (when prompts cross the provider's minimum size), Google, DeepSeek, xAI, and Alibaba inspect incoming requests for shared prefixes and cache them automatically. Vichar forwards the provider's cache metadata back to you in the response, and bills the cached portion at the model's `cached_input` rate.
For **Anthropic** and **AWS Bedrock Claude**, prompt caching is strictly opt-in via `cache_control` / `cachePoint` markers on the request body. To get automatic cache benefits without rewriting your requests, Vichar injects those markers for you on long system and user messages by default.
### Choosing a cache-write mode [#choosing-a-cache-write-mode]
**Project Settings → Caching → Provider Cache Writes** controls how the gateway treats those markers:
| Mode | Markers your client sends | Markers the gateway adds |
| ----------------------- | ------------------------- | ------------------------ |
| **Automatic** (default) | Forwarded | Added on long prompts |
| **Client-managed** | Forwarded | Never added |
| **Disabled** | Stripped | Never added |
**Automatic** suits requests that do not manage caching themselves. If you send long prompts sporadically — with gaps wider than the 5-minute TTL — you pay the cache-write premium (1.25× input for 5m, 2× for 1h) without ever benefiting from a cache read, so one of the other modes will be cheaper.
**Client-managed** hands the decision to each request: a request writes to the provider cache only if it carries its own markers. Use it when one API key serves both a coding tool that sets its own markers (Claude Code, Cursor, Cline) and other traffic that should not pay the write premium — Automatic would add markers to the latter, and Disabled would remove the former's.
**Disabled** turns provider caching off for the project entirely, including markers your client sends.
The mode applies to every upstream that takes an explicit cache marker, whichever field it uses for one — the gateway translates your `cache_control` into the marker the resolved provider expects. Providers that cache automatically with no marker at all are unaffected, since there is nothing to forward or strip. On models whose explicit caching is a request-level mode rather than a per-block marker, Client-managed and Disabled also send `prompt_cache_options: {"mode": "explicit"}`, so implicit caching does not write a cache the request never asked for. The [models page](https://app.vichar.io/dashboard) shows which models support prompt caching.
Where the upstream takes per-block markers, they are forwarded on `system` blocks, message text blocks, tool definitions, and `tool_result` blocks. A `cache_control` on an image or other non-text content block is not forwarded — put the breakpoint on an adjacent text block instead.
Changes saved through the dashboard or API take effect immediately — the gateway's cached project settings are invalidated on write.
To take advantage of automatic caching:
* Put stable content (system prompt, instructions, tool definitions, long documents) at the **start** of your messages
* Keep the variable portion (the latest user turn) at the **end**
* Reuse the same prefix across requests — even minor changes invalidate the cache
You can confirm the cache is working by inspecting `usage.prompt_tokens_details.cached_tokens` on the response. See [Cost Breakdown](https://docs.vichar.io/features/cost-breakdown) for the full list of usage fields.
```json
{
"usage": {
"prompt_tokens": 8200,
"completion_tokens": 150,
"prompt_tokens_details": {
"cached_tokens": 8000
},
"cost_details": {
"input_cost": 0.0006,
"cached_input_cost": 0.0008
}
}
}
```
In this example, 8,000 of the 8,200 prompt tokens were served from the provider's cache and billed at the cached rate.
### Pricing and routing [#pricing-and-routing]
Cached input tokens are billed at the model's published `cached_input` price (typically 10–25% of the regular input price, depending on the provider and model). Output tokens and any non-cached input tokens are billed at the normal rate.
When [Smart Routing](https://docs.vichar.io/features/routing) selects a provider for a large prompt (≥ 5,000 estimated tokens) or a new session, it estimates token costs from uncached input, cache reads, and output. It uses your project's observed token mix from the last 24 hours, falling back to workload defaults until enough usage exists. Session pricing applies even to a short opening prompt. Cache support alone receives no extra scoring preference by default. See [Prompt Caching and Token Costs](https://docs.vichar.io/features/routing#smart-routing-algorithm) for defaults, thresholds, and overrides.
## Explicit caching with `cache_control` [#explicit-caching-with-cache_control]
Some providers — most notably **Anthropic** — also support *explicit* cache control, where you mark specific content blocks as cacheable using a `cache_control` field. This gives you precise control over what gets cached and lets you opt into longer cache lifetimes than the default.
Explicit caching is provider-specific. Supported providers and TTLs at the time of writing:
| Provider | Models | Supported TTLs |
| -------------------- | ------------------------------ | -------------------- |
| Anthropic (Claude) | All Claude models | `5m` (default), `1h` |
| AWS Bedrock (Claude) | All Claude models | `5m` (default), `1h` |
| Alibaba (Qwen) | Qwen models with cache support | Provider-defined |
To mark content as cacheable, send the message content as an array of blocks and add a `cache_control` field to the block you want to cache:
```json
{
"model": "claude-haiku-4-5",
"messages": [
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are a helpful assistant. ",
"cache_control": { "type": "ephemeral", "ttl": "1h" }
}
]
},
{
"role": "user",
"content": "What is the capital of France?"
}
]
}
```
Use `ttl: "5m"` (the default if omitted) for short-lived caches that match a single user's session, and `ttl: "1h"` when the same prefix will be reused over a longer window (for example, a coding agent that keeps the same project context warm across many requests).
The same `cache_control` markers work on the native [`/v1/messages` endpoint](https://docs.vichar.io/features/anthropic-endpoint) (Anthropic request format), where cache usage comes back in Anthropic's own fields: `usage.cache_creation_input_tokens` (written this request) and `usage.cache_read_input_tokens` (served from cache).
### Minimum cacheable prompt length [#minimum-cacheable-prompt-length]
Every Claude model has a **minimum prompt length below which nothing is cached**. A `cache_control` marker on a shorter prompt is accepted without error, but the provider silently skips the cache write — the response reports `cache_creation_input_tokens: 0` and `cache_read_input_tokens: 0` on every call, no matter how often you repeat the request. This is Anthropic's documented behavior, not a dropped breakpoint; you would see exactly the same result calling Anthropic directly.
The threshold counts **all tokens up to and including the marked block** (tools, system, and preceding messages), and it varies by model.
The models endpoint (`GET https://api.vichar.io/v1/models`) exposes each model's exact threshold as `min_cacheable_tokens` on its provider entry, so clients can check it programmatically. Vichar's automatic marker injection uses the same threshold, which is why short system prompts never trigger automatic cache writes either.
### Verifying cache writes and the write premium [#verifying-cache-writes-and-the-write-premium]
The first request that creates a cache entry is billed at the provider's cache-write rate (1.25× input for `5m`, 2× for `1h`). On the OpenAI-compatible `/v1/chat/completions` endpoint, Vichar surfaces the write side in extended usage fields — OpenAI's standard format only has `cached_tokens` for reads, so these are gateway extensions:
```json
{
"usage": {
"prompt_tokens": 5232,
"prompt_tokens_details": {
"cached_tokens": 0,
"cache_write_tokens": 5222,
"cache_creation_tokens": 5222,
"cache_creation": {
"ephemeral_5m_input_tokens": 5222,
"ephemeral_1h_input_tokens": 0
}
},
"cost_details": {
"input_cost": 0.00003,
"cache_write_input_cost": 0.0065275,
"cached_input_cost": 0
}
}
}
```
`cache_write_tokens` (and its alias `cache_creation_tokens`) counts the prompt tokens written into the provider cache this request, `cache_creation` breaks the write down by TTL when both rates are in play, and `cost_details.cache_write_input_cost` is the exact USD amount billed at the write premium. A successful cache write on call 1 shows up as `cache_write_tokens > 0`; the matching read on call 2 shows up as `cached_tokens > 0` with `cached_input_cost` at the discounted rate.
### Mixing explicit markers with automatic injection [#mixing-explicit-markers-with-automatic-injection]
Anthropic requires cache breakpoints with longer TTLs to appear before shorter ones (blocks are processed in the order `tools`, `system`, `messages`). The markers Vichar injects automatically use the default 5-minute TTL, so they could never legally precede an explicit `ttl: "1h"` marker in your messages. To keep both features compatible:
* When your request contains an explicit `ttl: "1h"` marker in the **messages**, Vichar skips its automatic marker injection for that request entirely and forwards only your markers — the same behavior you would get calling the provider directly.
* A `ttl: "1h"` marker only on the **system** prompt does not disable automatic injection, since 5-minute breakpoints after it still satisfy the ordering rule.
* Explicit markers that use the default 5-minute TTL coexist with automatic injection (capped at 4 breakpoints total per Anthropic's limit).
This section describes the default **Automatic** mode. In **Client-managed** mode there is no injection to reconcile — your markers are forwarded exactly as sent, whatever their TTL.
Cache writes are billed at a premium (typically 1.25x for 5m and 2x for 1h on
Anthropic) the first time a cached block is created. After that, cache reads
cost roughly 10% of the regular input price. The break-even point is usually
one or two reuses — explicit caching is worth it whenever a marked block will
be sent more than once within its TTL.
Anthropic returns a per-TTL breakdown of cache writes when you mix `5m` and `1h` blocks:
```json
{
"usage": {
"cache_creation": {
"ephemeral_5m_input_tokens": 0,
"ephemeral_1h_input_tokens": 8000
},
"cache_read_input_tokens": 0
}
}
```
For providers that publish a separate explicit-cache read rate (for example, Alibaba Qwen charges 10% for explicit cache reads vs. 20% for automatic cache reads), Vichar detects the `cache_control` markers on your request and applies the explicit rate automatically.
## Related [#related]
* [Gateway Caching](https://docs.vichar.io/features/caching/gateway-caching) — serve identical requests entirely from Vichar at $0 cost
* [Caching Overview](https://docs.vichar.io/features/caching) — side-by-side comparison of provider caching vs. gateway caching
* [Cost Breakdown](https://docs.vichar.io/features/cost-breakdown) — full reference for the usage and cost fields on every response
* [Smart Routing](https://docs.vichar.io/features/routing) — how workload defaults, observed cache-hit rates, and output proportions influence provider selection