Docs
Point your toolsat LotusAI.
01 / Quick start
Two variables. That is the setup.
export ANTHROPIC_BASE_URL=https://api.lotusai.pro
export ANTHROPIC_AUTH_TOKEN=sk-lotus-...$env:ANTHROPIC_BASE_URL = "https://api.lotusai.pro"
$env:ANTHROPIC_AUTH_TOKEN = "sk-lotus-..."Get a key
Keys are issued by authorised partners. They start with sk-lotus-.
Set the variables
Base URL plus token, in your shell or the tool config.
Keep building
Streaming, tool use, vision and long context all pass straight through.
Base URL: https://api.lotusai.pro. Auth: send the key as x-api-key or Authorization: Bearer; both work on every endpoint. For OpenAI-format tools, add /v1 to the base URL.
02 / One-line installer
Or let a script do it.
curl -fsSL https://lotusai.pro/setup.sh | bashirm https://lotusai.pro/setup.ps1 | iex- Idempotent: a marked block in your profile is replaced, never appended twice.
- Pass the key non-interactively with LOTUSAI_API_KEY=sk-lotus-... in the environment.
- Writes ~/.claude/settings.json so Claude Code works from any terminal or editor.
- Prefer to read before you run? Both scripts are plain text at the URLs above.
03 / Set up your tool
Pick your editor.
Claude Code
Anthropic format- 1Run the one-line installer above, or set the two environment variables by hand.
- 2Or put them in ~/.claude/settings.json so they apply to every project.
- 3Restart Claude Code. Run /status inside it to confirm the base URL.
export ANTHROPIC_BASE_URL=https://api.lotusai.pro
export ANTHROPIC_AUTH_TOKEN=sk-lotus-...{
"env": {
"ANTHROPIC_BASE_URL": "https://api.lotusai.pro",
"ANTHROPIC_AUTH_TOKEN": "sk-lotus-..."
}
}Optional: pin a model with ANTHROPIC_MODEL, for example claude-fable-5-1[1m]. The [1m] suffix only tells Claude Code to size its own context budget at 1M; the API ignores it.
04 / API reference
Two formats, one key.
/v1/messages, OpenAI Chat Completions at /v1/chat/completions. Same key, same window, same models./v1/messagesAPI keyAnthropic Messages format, request and response. Supports system, multi-turn messages, image blocks, tools with tool_use and tool_result, and stream: true for server-sent events. Keys are sent as x-api-key; the anthropic-version header is accepted and optional.
curl https://api.lotusai.pro/v1/messages \
-H "x-api-key: sk-lotus-..." \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "Explain this regex: ^sk-lotus-" }]
}'curl -N https://api.lotusai.pro/v1/messages \
-H "x-api-key: sk-lotus-..." \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"stream": true,
"messages": [{ "role": "user", "content": "Stream me a limerick." }]
}'
# event: message_start
# event: content_block_start
# event: content_block_delta (repeats)
# event: content_block_stop
# event: message_delta
# event: message_stop{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"tools": [{
"name": "get_weather",
"description": "Current weather for a city",
"input_schema": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}],
"messages": [{ "role": "user", "content": "Weather in Lisbon?" }]
}
// The reply carries a tool_use block:
// { "type": "tool_use", "id": "toolu_...", "name": "get_weather", "input": { "city": "Lisbon" } }
// Run the tool, then send a tool_result block with the same id in the next user turn.POST /v1/messages/count_tokens is also available and takes the same body, returning { "input_tokens": n } without sending the request.
/v1/chat/completionsAPI keyOpenAI Chat Completions format. The request is translated to the Messages API and the reply is translated back, so messages with system roles, tools with tool_calls, and stream: true all behave as the OpenAI SDK expects. Use Authorization: Bearer for the key.
curl https://api.lotusai.pro/v1/chat/completions \
-H "Authorization: Bearer sk-lotus-..." \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-5",
"stream": false,
"messages": [
{ "role": "system", "content": "You are a terse senior engineer." },
{ "role": "user", "content": "Why is my build slow?" }
]
}'/v1/modelsNo authLists every model id with its display_name and context_window. Public, so tools can populate their model pickers before a key is entered. GET /v1/models/:id returns one entry.
{
"data": [
{
"id": "claude-fable-5-1",
"display_name": "Claude Fable 5.1",
"type": "model",
"context_window": 1000000
},
{
"id": "claude-fable-5",
"display_name": "Claude Fable 5",
"type": "model",
"context_window": 1000000
},
{
"id": "claude-opus-5",
"display_name": "Claude Opus 5",
"type": "model",
"context_window": 1000000
}
],
"has_more": false
}/api/key-status?key=sk-lotus-...No authEverything the Check key page shows, as JSON: plan, active and expiry flags, window usage as a percentage, per-family split, reset time, request counts and the last twenty requests. Token budgets are never exposed as raw numbers. An unknown key returns { "status": "error", "error": "API key not found" }.
{
"status": "found",
"name": "Team key",
"keyPrefix": "sk-lotus-a1b2",
"isActive": true,
"planName": "Studio",
"rateLimit": 60,
"unlimited": false,
"usagePercent": 42.5,
"byModel": { "fable": { "percent": 10.2 }, "opus": { "percent": 0 }, "sonnet": { "percent": 30.1 }, "haiku": { "percent": 2.2 } },
"windowActive": true,
"windowStartedAt": "2026-10-03T08:10:00.000Z",
"windowResetAt": "2026-10-03T13:10:00.000Z",
"expiresAt": "2026-11-01T00:00:00.000Z",
"lastUsedAt": "2026-10-03T09:42:11.000Z",
"createdAt": "2026-09-01T12:00:00.000Z",
"isExpired": false,
"isOverLimit": false,
"totalRequests": 1842,
"last24h": { "requests": 96 },
"recentLogs": [{ "model": "claude-sonnet-5", "status": 200, "time": "2026-10-03T09:42:11.000Z" }]
}05 / Models
Every model, one key.
model field. This list is built from GET /v1/models and refreshes every five minutes.| Model | Best for | Context |
|---|---|---|
Fable 5.1new claude-fable-5-1The most capable model on the list: demanding reasoning and long-horizon agentic work. Knows the world up to June 2026, the most recent cutoff of anything we serve. | The most capable model on the list: demanding reasoning and long-horizon agentic work. Knows the world up to June 2026, the most recent cutoff of anything we serve. | 1M |
Fable 5 claude-fable-5The Fable before 5.1. Still holds its own on long agentic runs and whole-codebase work. Keep it pinned if your prompts are tuned to it. | The Fable before 5.1. Still holds its own on long agentic runs and whole-codebase work. Keep it pinned if your prompts are tuned to it. | 1M |
Opus 5new claude-opus-5The newest Opus. Thinks before it answers by default, and knows the world up to May 2026. Point it at the hard stuff. | The newest Opus. Thinks before it answers by default, and knows the world up to May 2026. Point it at the hard stuff. | 1M |
Sonnet 5new claude-sonnet-5Frontier-class intelligence at Sonnet speed. The new default for everyday building. | Frontier-class intelligence at Sonnet speed. The new default for everyday building. | 1M |
Opus 4.8 claude-opus-4-8Adaptive thinking and agentic coding. Whole-repo refactors that need it to actually hold the plot. | Adaptive thinking and agentic coding. Whole-repo refactors that need it to actually hold the plot. | 1M |
Opus 4.7 claude-opus-4-7Still a monster. Keep pinning it if your prompts are tuned to it. | Still a monster. Keep pinning it if your prompts are tuned to it. | 1M |
Opus 4.6 claude-opus-4-6Deep reasoning on big, gnarly problems. | Deep reasoning on big, gnarly problems. | 1M |
Sonnet 4.6 claude-sonnet-4-6The everyday sweet spot: most of what you build, most days. | The everyday sweet spot: most of what you build, most days. | 1M |
Opus 4.5 claude-opus-4-5A known quantity for pipelines you don’t want to re-tune. | A known quantity for pipelines you don’t want to re-tune. | 200K |
Sonnet 4.5 claude-sonnet-4-5-20250929Solid general-purpose work when you want the older behaviour. | Solid general-purpose work when you want the older behaviour. | 200K |
Haiku 4.5 claude-haiku-4-5-20251001Autocomplete, quick edits, classification, anything high-volume where latency is the whole game. | Autocomplete, quick edits, classification, anything high-volume where latency is the whole game. | 200K |
Opus 4.1 claude-opus-4-1-20250805Pinned for reproducibility. | Pinned for reproducibility. | 200K |
Opus 4 claude-opus-4-20250514Pinned for reproducibility. | Pinned for reproducibility. | 200K |
Sonnet 4 claude-sonnet-4-20250514Pinned for reproducibility. | Pinned for reproducibility. | 200K |
06 / The 5-hour window
Usage that refills.
How the clock works
- Starts on first use. The window opens with the first request that consumes tokens, not at a fixed hour.
- Runs for five hours. Every request in that span counts towards the same budget.
- Then resets completely. Usage goes back to zero and the next request opens a new window.
What counts
Input and output tokens of each request, weighted per model family. Prompt-cache reads do not count. Larger models draw the budget down faster than smaller ones.
What happens at the limit
Requests are answered with a short message stating the time until the window refills, so your tool keeps running instead of crashing on an error. Requests resume the moment the window resets.
Where to watch it
Check key shows the gauge, the per-model split and a live countdown. The same data is at GET /api/key-status.
07 / Rate limits and status codes
What the gateway returns.
Per key
req / min
Each key has a requests-per-minute ceiling set by the partner. It is shown on the Check key page. Exceeding it returns 429.
Per IP
200 / min
A global guard against floods, counted per source address across all endpoints. Normal editor use never approaches it.
Per window
5 hours
The token budget described above. Reaching it does not return an error; see the window section.
| Code | error.type | Meaning |
|---|---|---|
| 200 | Success. Also used for the limit-reached, expired and disabled notices, which arrive as a normal assistant message so tools display them. | |
| 400 | invalid_request_error | Malformed body, empty messages array, unknown parameter, or a request too large to process. |
| 401 | authentication_error | Missing or unrecognised key. |
| 404 | not_found_error | Unknown model id or path. |
| 413 | invalid_request_error | Payload exceeds the size limit. Shrink images or trim history. |
| 429 | rate_limit_error | Per-key or per-IP request rate exceeded. Honour the Retry-After header and back off. |
| 500 / 502 | api_error | Something failed on our side or between us and the model. Retry with backoff. |
| 503 | maintenance | Planned maintenance, with Retry-After in seconds. |
| 529 | overloaded_error | Model capacity is saturated. Retry after a short pause. |
{
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "Rate limit reached. Please try again shortly."
}
}08 / Troubleshooting
When something looks off.
401 authentication_error+
sk-lotus-, and confirm it is sent as x-api-key or Authorization: Bearer. Then paste it into Check key.The reply says the usage limit was reached+
Model not found+
claude-sonnet-5. GET /v1/models returns the live list.Cursor says the endpoint failed verification+
/v1 suffix: https://api.lotusai.pro/v1. Add the model id as a custom model after verifying.Claude Code still goes to the default endpoint+
~/.claude/settings.json (the installer does this), then restart Claude Code and run /status.Context window exceeded+
Changes to a config file do nothing+
503 with Retry-After+
Ready
Have a key? Go build.
Still need one? Keys are issued by authorised partners, and we will put you in touch with one.