Get API key

getdeepseekapi.comDocs

DeepSeek API: First request in five minutes

Get your first response in five minutes by pointing your existing OpenAI-compatible client to our base URL and authenticating with the key generated during signup.

Base URL and Authentication

Our API follows the standard OpenAI interface, so you only need to change two values in your client configuration. Set the base URL to https://api.getdeepseekapi.com/v1 and use the API key displayed immediately after you create your account on the Get API key page. No phone number or credit card is required to start, and you receive $0.50 in trial credit that lasts for seven days.

Each account allows only one active key at a time. If you rotate your key, the previous one is immediately revoked. This simplifies secret management across development and production environments without needing complex permission scopes.

Send a Chat Completion

Make a standard POST request to /v1/chat/completions to receive a text response. The endpoint accepts a messages array containing your system prompt and user query, along with a model identifier set to uncensored. This model runs on our own GPU servers and is tuned to answer without content refusals for lawful adult use.

Ensure your request body stays under 8 MB. If the key is invalid or expired, the server returns a 401 error. If your prepaid credit is exhausted, you receive a 402 error requiring a top-up. You can add funds starting at $10 via crypto (USDT or USDC).

curl https://api.getdeepseekapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Install and Use the Python SDK

Using the official OpenAI Python library is the fastest way to integrate. Install the package via pip, then configure the client with your custom base URL and API key. The model parameter must be set to uncensored to ensure you are talking to our specific instance.

This approach works with any OpenAI-compatible SDK, including those used for Node.js or Go. The client handles serialization and retries automatically. Remember that the model id is not a vendor name like GPT-4; it is a specific identifier for our uncensored model.

from openai import OpenAI

client = OpenAI(base_url="https://api.getdeepseekapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js Integration

For JavaScript environments, install the OpenAI SDK and configure it similarly to the Python example. Set the baseUrl to our endpoint and provide your API key. The Node SDK supports both synchronous and asynchronous patterns.

When constructing the request, specify model: 'uncensored' in the parameters. This ensures the client sends requests to the correct model endpoint. The SDK handles streaming and non-streaming responses uniformly, allowing you to swap implementations without changing the core request logic.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.getdeepseekapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Enable Streaming (SSE)

Set the stream parameter to true to receive a Server-Sent Events (SSE) response. The server sends a series of chunks, each containing a portion of the generated text. This reduces perceived latency for long responses and allows the client to display tokens as they are produced.

Ensure your client code properly handles the stream termination event. Streaming does not change the pricing or token counting; you are still charged per 1M tokens for input and output. The total context window remains limited to 100,000 tokens for the combined prompt and completion.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Limits, Errors, and Context Window

Your API key is limited to 300 requests per minute. If you exceed this threshold, the server returns a 429 error indicating a rate limit. Each request body must not exceed 8 MB. The context window supports 100,000 tokens for the sum of the prompt and the completion, meaning you must account for both input and output lengths when designing your application.

Common errors include 401 for invalid keys, 402 for insufficient prepaid credit, and 429 for rate limits. You can regenerate your API key at any time from your account settings, which immediately invalidates the old key. This is useful if you suspect a leak or need to rotate credentials between environments.

API specifications

Before you integrate, here is exactly what you get with a key.

ItemValue
ProtocolOpenAI Chat Completions schema; official openai SDKs work unchanged
EndpointsPOST /v1/chat/completions · GET /v1/models
AuthenticationBearer token in the Authorization header
Model IDuncensored
Base URLhttps://api.getdeepseekapi.com/v1
Context window100,000 tokens (prompt + completion together)
Sampling parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
JSON modeJSON object mode via response_format json_object
SSE streamingSupported (stream: true), usage included at the end
Max output16,000 tokens max; 2,048 if max_tokens is not set
Tools / tool callsYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Rate limit300/min per key
Max body8 MB request body
Concurrency8 requests at the same time per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Subscriptionno monthly fee; paid credit does not expire
Top-upcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Trial credit$0.50 for 7 days, no card
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
How you payprepaid credit, charged by real token usage; errors and refusals are free
Volume bonus+5% on $50+, +10% on $100+
Sign-insign in with Google or with e-mail + password
Content policyuncensored for adults; the only hard rule: no sexual content involving minors
Keysone active key per account; a new key replaces the old one

Error codes

The type field is stable, the message is for humans. Errors cost nothing.

StatusTypeReason
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditbalance is empty — top up, requests resume at once
403content_blockedsexual content involving minors — refused, not billed
404not_foundunknown endpoint
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

What is the context window size?

The context window is 100,000 tokens, calculated as the sum of the input prompt tokens and the output completion tokens. It is not 100,000 tokens available exclusively for input.

Does this service support tool calling?

Yes, the <code>/v1/chat/completions</code> endpoint supports tool and function calling. You can define tools in your request and the model will return structured arguments for them.

How do I handle a 402 error?

A 402 error indicates that your prepaid credit balance is zero. You can top up your account starting at $10 using crypto (USDT or USDC). Credits never expire.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key