HomeDocumentation
API Quick Integration Guide
Integrate our OpenAI-compatible API in minutes to get direct, uncensored responses. This guide covers authentication, main endpoints, streaming, and rate limit management so you can start using your uncensored AI immediately.
Authentication and Base URL
To consume our AI API, you will need an API key and the correct base URL. Sign in on the Get API key page with your email and password to get your key immediately; no card is required for the trial. Use the base URL https://api.llmsincensura.com/v1 with any OpenAI-compatible client. Authentication is done by sending your key in the Authorization: Bearer YOUR_API_KEY header. This standard method ensures your uncensored LLM works with the tools you already know.
The model you will serve is uncensored. Make sure to configure it in your requests. There are no complex routes; you send text and receive raw text. If your application requires a specific endpoint, this is the only one you need for chat. Simplicity is key for fast, frictionless integration.
Endpoint /v1/chat/completions
The core of the uncensored AI API is the chat completion endpoint. Send a list of messages (system, user, assistant) and receive a complete response. This direct approach avoids the content refusals common in other models for adult or controversial topics. The JSON structure is identical to OpenAI's, making it easy to switch if you come from another platform.
Remember that we only support this endpoint for text. There are no embeddings or image generation. To get started, you can test a simple request with curl to verify that your key works and that the response arrives unfiltered.
curl https://api.llmsincensura.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python Integration
Python developers can use the official openai library by adapting the base URL. This allows you to leverage the existing ecosystem of tools and SDKs. Configure the key and base_url to point to our servers. The uncensored model will process your input maintaining context up to 100,000 tokens.
This configuration is ideal for automation scripts or backend applications that need fast, uncensored responses. Response speed is usually high due to the dedicated infrastructure on our own GPU servers.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmsincensura.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node.js Integration
In the JavaScript/Node ecosystem, integration is equally straightforward. Use the OpenAI SDK and override the base URL. This allows web applications or Node servers to consume unrestricted AI without changing your existing code logic. Error handling and promises work exactly as you expect with other OpenAI clients.
It is important to remember that the model is not GPT or Claude, but an open-weight model fine-tuned specifically for this service. This ensures consistent, uncensored behavior. Configure environment variables to keep your key secure in production.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmsincensura.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
SSE Streaming
To improve the user experience, enable streaming by setting stream: true in your request. You will receive text chunks as they are generated, reducing perceived latency. This is crucial for real-time chat applications where users expect to see the response letter by letter.
Streaming maintains the same non-filtering properties as complete responses. You can process these chunks on the client or server side to render the response progressively. It is the best way to showcase the power of an uncensored AI chat in interactive applications.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits and Error Handling
Your key has a rate limit of 300 requests per minute. If you exceed this threshold, you will receive a 429 error. There is also an 8 MB request body limit. If your key is invalid, you will get a 401; if you have no credit, a 402. The maximum context is 100k tokens combined (prompt + response).
The only strict content restriction is the prohibition of sexual content with minors, which always applies. For other topics, the model is very permissive. Check the standard OpenAI error codes to debug your integration quickly.
Capabilities and limits
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Item | Value |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.llmsincensura.com/v1 |
| API key | Bearer token in the Authorization header |
| Model | uncensored |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Max output | prompt + completion fit within 100,000 tokens; max_tokens optional, no separate output cap |
| JSON mode | JSON object mode via response_format json_object |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Max context | 100,000 tokens (prompt + completion together) |
| Requests per minute | 300/min per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Parallel requests | 8 requests at the same time per key |
| Max body | 8 MB request body |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Free trial | $0.50 for 7 days, no card · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up |
| Payment | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Bonus credit | +5% on $50+, +10% on $100+ |
| Credit expiry | no monthly fee; paid credit does not expire |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Keys | one active key per account; a new key replaces the old one |
| Content | adult content allowed; sexual content involving minors is refused |
| Sign-in | sign in with Google or with e-mail + password |
Errors and what to do
Every error is JSON with a type you can switch on. You are never charged for an error.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Frequently asked questions
What happens if I run out of balance?
If your balance reaches zero, you will receive a 402 error when making requests. You can top up from $10 with cryptocurrency (USDT or USDC) and the credit never expires. No mandatory monthly subscriptions.
Can I use this API for commercial use?
Yes, the model is designed for general use, including commercial. There are no industry-based usage restrictions, as long as the content is not sexual with minors. Your privacy is protected and we do not use your prompts to train other models.
How do I regenerate my API key?
You can regenerate your key at any time from your account dashboard. When you generate a new one, the previous key stops working immediately. This is useful if you think your key has been compromised or if you need different keys for development and production environments.
Your key is one step away
Create an account, copy the key, and change the base URL. That's it.
Get API key