Switch your client in three lines
Migrate your client to our uncensored API in minutes by updating the base URL and authentication header. This guide covers the essential endpoints, streaming, and constraints for the OpenAI-compatible chat-completions interface.
Base URL and Authentication
To use the API, point your client to the base URL https://api.mistralapi.cc/v1. You do not need a complex setup or SDK installation for basic requests; standard HTTP clients work. Authentication relies on a Bearer token passed in the Authorization header. Your key is issued immediately upon signup via Google or email, with no credit card required for the trial.
The service runs a single uncensored model identified as uncensored. It is not GPT, Claude, Grok, or any other vendor's model. It is an open-weight model tuned to answer without refusals for lawful adult use. When switching from another provider, simply update the base URL and API key in your configuration. All requests are charged against prepaid token credit. Errors and refusals do not consume balance.
Send a Chat Completion
The primary endpoint is POST /v1/chat/completions. Send a JSON body containing the model set to uncensored, along with your messages array. The API returns text output. No embeddings, images, or audio are generated. The request body is limited to 8 MB. If you exceed this limit, the request fails immediately without charging credit.
You can control generation via parameters like temperature, top_p, stop, and seed. The context window supports 64,000 tokens total, with a maximum output of 16,000 tokens per request. If you do not specify max_tokens, the output is capped at 2,048 tokens.
curl https://api.mistralapi.cc/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK
Use the official openai Python package. Configure the client with the custom base URL and your API key. This approach works with any code written for OpenAI-compatible endpoints. The model ID remains uncensored.
Streaming and non-streaming requests behave identically to standard implementations. Token usage is reported in the final chunk of the response. Function calling and JSON mode are supported via the response_format and tools parameters.
from openai import OpenAI
client = OpenAI(base_url="https://api.mistralapi.cc/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node SDK
The Node.js SDK follows the same pattern. Initialize the client with the provider's base URL overridden to https://api.mistralapi.cc/v1. Pass the API key in the header. This allows you to reuse existing client code from other providers with minimal changes.
Ensure you handle the choices[0].message.content field for text output. The API supports standard parameters like presence_penalty and frequency_penalty. No fine-tuning or model routing is available; you are connected directly to the uncensored model instance.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.mistralapi.cc/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
Enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream. Each chunk contains a partial delta of the response. Token usage details are included in the last chunk of the stream.
Streaming reduces perceived latency for long outputs. The server processes the full context window before sending the final token count. Errors during streaming are returned as standard HTTP error codes, not within the stream body.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Rate Limits and Constraints
Each API key is limited to 300 requests per minute and 8 concurrent requests. The request body size is capped at 8 MB. If you exceed these limits, the API returns a 429 Too Many Requests error. These errors do not consume credit.
Authentication failures return a 401 Unauthorized error. Insufficient funds return a 402 Payment Required error. Credit is prepaid and never expires. You can top up using USDT (TRC20) or USDC (Base). The trial credit of $0.50 is valid for 7 days. No card is needed for the trial.
Questions and answers
Does the API support image or audio generation?
No. The API serves text output only via the chat-completions endpoint. There are no embeddings, images, audio, or video generation features.
What happens if I run out of credit?
Requests return a 402 error. No tokens are charged for the failed request. You must top up via crypto to resume service. Credit never expires.
Is the model one of the major vendors like GPT or Claude?
No. The model is identified as <code>uncensored</code>. It is an open-weight model run on our servers, tuned for fewer refusals, but it is not GPT, Claude, Grok, DeepSeek, or any other vendor's model.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.