HomeDocs
Quickstart for Claude Code users
Connect Claude Code or any OpenAI-compatible agent to our proxy in under two minutes. This quickstart covers the base URL, authentication, and the exact endpoints you need for coding tasks.
Authentication
Every request to our proxy requires a valid API key. You receive this key immediately after signing up on the Get API key page. Include it in the Authorization header as a Bearer token. If the key is missing or invalid, the API returns a 401 error. You can regenerate your key at any time from your dashboard, which instantly revokes the previous one. This ensures that lost or compromised keys do not allow continued access to your prepaid credits.
Chat Completions Endpoint
Send your coding prompts to the standard chat completions endpoint. The base URL for all requests is https://api.claudecodeapikey.com/v1. Use the model ID uncensored to access our large language model. This model is tuned to answer without content refusals for lawful adult use, making it ideal for generating code without unnecessary blocks. The endpoint supports both standard and streaming responses. Below is an example of a basic request using curl.
curl https://api.claudecodeapikey.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
The request body must stay under 8 MB. If you exceed this limit, the server will reject the payload. The model handles a context window of 100,000 tokens, combining both input and output tokens. This allows for substantial code snippets and long conversation histories without immediate truncation.
Python SDK
For developers using Python, you can integrate our API using the official OpenAI SDK. Simply point the client to our base URL and provide your API key. The uncensored model ID works exactly like other OpenAI-compatible models in the SDK. You can call the completion method and handle the response as a standard text object. This approach is ideal for batch processing or integrating into larger automation scripts.
from openai import OpenAI
client = OpenAI(base_url="https://api.claudecodeapikey.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Remember that the SDK handles the JSON serialization for you. You only need to ensure that your base_url variable points to our proxy. The response object contains the generated text, which you can then write to your files or feed back into the conversation context.
Node SDK
Node.js developers can use the OpenAI npm package to interact with our proxy. Configure the client with the custom base URL and your API key. The usage pattern mirrors the Python implementation: create a client, call the chat completions endpoint, and process the result. This is useful for server-side code generation or real-time coding assistants.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.claudecodeapikey.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
The Node SDK manages the HTTP connections efficiently. Ensure that you are using a version of the SDK that supports custom base URLs. The response structure remains consistent with the OpenAI standard, making it easy to swap between different OpenAI-compatible providers if needed.
Streaming Responses
For real-time coding assistance, enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream. Each chunk contains a partial response, allowing you to display code as it is generated. This significantly improves the user experience in interactive coding agents.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Handle the stream events in your code to accumulate the final response. Streaming does not consume more tokens than a non-streaming request; it only changes how the data is delivered. This is particularly useful for long code blocks where you want to show progress immediately.
Rate Limits & Quotas
Our API enforces a limit of 300 requests per minute per API key. If you exceed this, you will receive a 429 error. You should implement retry logic with exponential backoff in your applications. Additionally, each account is limited to one API key, which can be regenerated if needed. The 402 error indicates that your prepaid credit has been exhausted. You can top up from $10 using crypto (USDT or USDC). Credits never expire, so you can pause and resume usage at your own pace without losing your balance.
Technical specifications
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Parameter | Details |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Authentication | Authorization: Bearer YOUR_KEY |
| Model ID | uncensored |
| Base URL | https://api.claudecodeapikey.com/v1 |
| Streaming | Supported (stream: true), usage included at the end |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Context window | 100,000 tokens, input and output combined |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Structured output | response_format: {"type": "json_object"} |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Rate limit | 300 requests per minute per key |
| Concurrency | 8 requests at the same time per key |
| Max body | up to 8 MB per request |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Trial credit | $0.50 for 7 days, no card |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Bonus credit | +5% from $50, +10% from $100 |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Credit expiry | no monthly fee; paid credit does not expire |
| Content | adult content allowed; sexual content involving minors is refused |
| Key management | one key per account, regenerate any time (the old one stops working) |
| Account | sign in with Google or with e-mail + password |
Error reference
The type field is stable, the message is for humans. Errors cost nothing.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Is this the official Anthropic API?
No, this is an independent proxy service. We host our own uncensored large language model that is compatible with the OpenAI chat completions format. It is not GPT, Claude, or any other vendor's model.
What happens if I run out of credits?
Requests will return a <code>402</code> error. You can top up your account at any time from $10 using crypto (USDT or USDC). Your unused credits never expire, so you can add funds whenever convenient.
Does the model refuse content?
The model is tuned to answer without refusals for lawful adult, fictional, or controversial topics. The only hard limit is that sexual content involving minors is always blocked.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.