HomeGuide
How to Integrate Uncensored AI into Your Application
Updated LLM Sin Censura
Integrating an uncensored AI into your application allows you to get direct responses from the model without the frequent refusals of major providers. This technical guide explains how to connect an OpenAI-compatible endpoint, manage context, and handle streaming to build smooth, uncensored user experiences.
What is Uncensored AI?
An uncensored AI refers to a large language model (LLM) that has been fine-tuned or configured to not reject adult content, controversies, or niche topics based on rigid corporate policies. Unlike standard commercial models that may return 400 errors or generic text when they detect certain patterns, an uncensored model returns exactly what is requested, as long as it is legal.
This feature is crucial for creativity, research applications, or environments where the model's freedom of expression is a priority. It does not mean the model is 'dumb' or lacks intelligence; it simply removes the layers of superficial moderation that sometimes block valid responses. For the developer, this simplifies error handling logic, as you do not need to predict what content will be rejected by an external filter.
- Predictability: You get the full text without arbitrary cuts.
- Control: You decide what content to display, not the API.
- Efficiency: Fewer failed API calls due to incorrect moderation.
Initial SDK Setup
To integrate our uncensored AI, we will use the standard OpenAI SDK, as our endpoint is 100% compatible with its protocol. This allows you to use libraries you already know, such as openai in Python or openai in Node.js. You only need to adjust the base URL and provide your API key.
The advantage of this approach is portability. If you decide to switch to another provider that also offers a compatible endpoint in the future, the change in your code will be minimal. Here is how to configure the client in Python pointing to our service:
from openai import OpenAI
client = OpenAI(base_url="https://api.llmsincensura.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Remember that compatibility means we use the same JSON request and response format, which facilitates migration from other platforms or the development of unit tests using standard mocks.
API Key Authentication
Security in our API is managed via a single API key per account. This key is generated immediately after registration and shown once. It is important to understand that there are no multiple keys by default; your account has a primary identifier for your requests.
If you need to revoke access or generate a new token for security, you can regenerate your key in your account settings. Note that this invalidates the previous key instantly. It is not a way to increase your rate limit, but a security measure to have only one active control point. Store your key in environment variables and never expose it in client-side code if your application is public.
- Generation: Shown upon registration.
- Regeneration: Invalidates the previous key.
- Usage: Sent in the
Authorization: Bearer YOUR_KEYheader.
Sending the First Request
The basic operation is a POST request to /v1/chat/completions. You must specify the model as uncensored, define the system prompt if necessary to set the tone, and provide the message history or the user's simple query. The model will process the input and return a complete response in JSON format.
This is the simplest flow to validate that your integration is working. Ensure that your application correctly handles the model's responses, which include the generated content and usage metadata, such as the number of tokens consumed.
curl https://api.llmsincensura.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'This request is the basis for any chat, content generation, or text analysis application. If the model responds correctly, your network and authentication configuration is correct.
Handling streaming
For interactive applications like chatbots, streaming is essential. It allows tokens to be sent to the client as they are generated, improving the perception of latency. Our API supports Server-Sent Events (SSE) for this purpose.
By enabling streaming in your SDK, you will receive multiple partial events. You must concatenate these fragments to build the complete response. This is especially useful when the model generates long texts, as the user starts seeing text almost instantly.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Streaming does not consume more tokens than a non-streaming request; it only changes how the data is transported. It is the recommended practice for any modern user interface that relies on AI.
Managing the 100k Context Window
One of the biggest challenges in LLM development is the context limit. Our API offers a context window of 100,000 tokens. This is enough to load extensive documents, long conversation histories, or large knowledge bases in a single request.
To leverage this efficiently, you must manage the sliding window in your application. As the conversation grows, you can remove the oldest messages from the history you send to the model to stay within the limit. The context includes both input tokens (prompt) and output tokens (completion).
- Capacity: 100k total tokens.
- Strategy: Implement a circular buffer for messages.
- Cost: Tokens are billed by usage, so optimize context usage.
Tool Calling Usage
Function calling allows the model to interact with your external code. You can define functions like get_weather or search_database and the model will return a structured argument to call them. This is essential for applications that need real-time data or specific actions.
Our uncensored AI supports function calling in the same way as the OpenAI standard. You define the tools in the tools parameter and the model responds with a call ID. Your code executes the function and then sends the result back to the model to generate the final response based on that data.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmsincensura.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);This approach turns a passive language model into an active agent capable of performing complex tasks in your application.
Limits and Quotas
To ensure service stability, our API has rate limits and size limits. Each API key is limited to 300 requests per minute. This is sufficient for most medium-use applications, but if your application has traffic spikes, you should consider implementing a request queue on your side.
Additionally, the request body must not exceed 8 MB. This limits the size of text or JSON files you can send in a single request. If you need to send more data, consider splitting it into multiple requests or summarizing it before sending.
| Resource | Limit |
|---|---|
| Speed | 300 req/min |
| Request Size | 8 MB |
| Model | uncensored |
Billing and Top Ups
The pricing model is pay-as-you-go, with no mandatory monthly subscriptions. You are charged per million tokens. The cost is $0.25 per 1M input tokens and $1.00 per 1M output tokens. This structure is transparent and predictable.
You can top up your account with a minimum of $10 using cryptocurrencies (USDT or USDC). If you top up larger amounts, you receive bonuses: an extra 5% for top-ups of $50 or more, and an extra 10% for top-ups of $100 or more. The credit never expires, so you can use it when you need it without time pressure.
- Input Price: $0.25 / 1M tokens.
- Output Price: $1.00 / 1M tokens.
- Bonuses: +5% starting from $50, +10% starting from $100.
Frequently Asked Questions
Is this model a copy of GPT or Claude?
No. It is an open-weight model running on our own GPU servers, fine-tuned specifically to be uncensored. It is not GPT, Claude, Gemini, or any other external provider model, although it is compatible with their API format.
What happens if I exceed the 300 requests per minute limit?
The API will return a rate limit error. You will need to wait for the counter to reset or implement exponential backoff retry logic in your application. Regenerating your key does not increase this limit, as it is per account and key.
Are my prompts used to train the model?
No. Our privacy policy is strict. The prompts you send to the API are not used for model training, ensuring the confidentiality of your data and conversations.
Can I use this API for NSFW content?
Yes, the model is designed to not censor legal adult content. However, it always blocks sexual content involving minors, regardless of jurisdiction. For the rest of adult content, the model does not offer refusals based on convenience.
Your key is one step away
Create an account, copy the key and change the base URL. That's it.
Get API key